CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual Entailment
Haoyu SongLi DongWeinan ZhangTing LiuFuru Wei
Demonstrates that pretrained CLIP models can serve as strong zero- and few-shot vision-language learners on visual question answering and visual entailment through generative question-to-statement prompt conversion and parameter-efficient bias-and-normalization tuning.
Developing artificial intelligence systems that simultaneously understand images and text typically demands massive datasets with expensive human annotations. Pre-trained vision-language models such as Contrastive Language-Image Pretraining (CLIP) learn broad visual representations from hundreds of millions of web image-caption pairs, but directly adapting them to complex reasoning tasks like visual question answering (VQA) has historically yielded poor, near-chance results. This limitation creates a major bottleneck for organizations aiming to deploy vision-language capabilities without incurring large data-labeling and training expenses.
The article demonstrates how to effectively transfer CLIP’s capabilities to vision-language understanding tasks under zero-shot and few-shot conditions without any additional large-scale pre-training. To achieve this, the authors evaluated CLIP on two core tasks: visual question answering (using the VQAv2 benchmark) and visual entailment (using the SNLI-VE benchmark). The approach introduces a two-step framework called TAP-C (Template-Answer-Prompt then CLIP discrimination), which uses a generative language model (T5) combined with rule-based dependency parsing to convert questions into fill-in-the-blank statements and filter out implausible answers. For few-shot adaptation, the authors developed BiNor, a parameter-efficient fine-tuning strategy that updates only the bias and normalization parameters (fewer than 0.3% of the total model weights).
The findings show that TAP-C dramatically improves CLIP’s zero-shot visual question answering performance from a prior baseline of about 21–23% to nearly 39%, outperforming larger specialized zero-shot baselines like Frozen (29.50%). In visual entailment, the article demonstrated a strong cross-modality transfer: training a lightweight classifier solely on text-text premise pairs allowed CLIP to evaluate image-text pairs with over 64–67% accuracy, nearly matching supervised text-only performance. Furthermore, when fine-tuning on limited labeled examples (few-shot), the BiNor method improved VQA accuracy to approximately 50%, consistently outperforming full model fine-tuning and bias-only tuning while avoiding overfitting.
These results demonstrate that organizations can repurpose web-pretrained vision-language models for specialized multimodal tasks without massive labeled datasets or compute-heavy retraining pipelines. Fine-tuning less than 0.3% of network parameters provides a fast, low-cost path to deploying capable vision-language systems. However, decision-makers should note clear operational limitations: CLIP models continue to struggle with fine-grained object counting in localized areas and often fail to differentiate subtle spatial relationships (such as distinguishing actions in the background versus the foreground). Before deploying these methods into high-stakes operational settings, teams should pilot the pipeline on domain-specific data and explore integrating stronger text encoders to mitigate spatial and semantic reasoning errors.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. CLIP’s image–text contrastive pretraining and zero-shot transfer provide the foundation this paper tests for few-shot VQA and visual entailment.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). CoOp introduces few-shot, parameter-efficient adaptation of frozen CLIP, providing a useful precursor for understanding how this paper adapts CLIP with limited labels.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). CLIP-Adapter demonstrates lightweight few-shot adaptation of frozen CLIP, clarifying the parameter-efficient fine-tuning approach this paper investigates.
No sufficiently relevant recommendations were found.
