Revisiting the Role of Language Priors in Vision-Language Models
Zhiqiu LinXinyue ChenDeepak PathakPengchuan ZhangDeva Ramanan
Introduces VisualGPTScore to repurpose generative vision-language models for zero-shot retrieval tasks and provides a training-free debiasing method that corrects for linguistic priors to achieve state-of-the-art accuracy across vision-language benchmarks.
Vision-language models are widely adopted for multimodal understanding tasks because they can operate out-of-the-box without task-specific fine-tuning. However, standard contrastive models often struggle with complex compositional reasoning—such as distinguishing "the horse is eating the grass" from "the grass is eating the horse"—while existing evaluation benchmarks frequently suffer from hidden linguistic biases that obscure true visual understanding.
The article aims to evaluate whether generative vision-language models can be repurposed for matching tasks and demonstrates a probabilistic method to control for language bias without model retraining.
The authors analyze image-conditioned generative models across nine retrieval and alignment benchmarks, including ARO, SugarCrepe, VL-CheckList, COCO, Flickr30K, and Winoground. They introduce the Visual Generative Pre-Training Score, which uses the model's conditional probability of generating a text caption given an image as a retrieval matching score. To resolve mismatches between pre-training language distributions and test distributions, they introduce a training-free debiasing formula that adjusts match scores using an estimated language prior derived via efficient Monte Carlo sampling of Gaussian noise images.
The analysis reveals four central findings. First, off-the-shelf generative scoring consistently outperforms traditional discriminative models and heavily engineered baselines across compositionality benchmarks without requiring additional training data. Second, multiple widely used benchmarks contain unnatural negative captions that can be solved by "blind" language-only models without looking at images, achieving up to 98–99% accuracy on benchmarks like ARO. Third, on realistic benchmarks where negative captions are plausible, raw generative models can wrongly favor common phrases; applying the proposed debiasing technique resolves this bias and improves accuracy, for example boosting Winoground text scores from 27.5% to 36.6% and ImageNet zero-shot classification from 18.6% to 40.0%. Fourth, applying debiased generative scoring to modern vision-language models like LLaVA-1.5 sets a new state-of-the-art on demanding image-text alignment tasks, outperforming complex pipelines that rely on auxiliary models like ChatGPT.
These findings indicate that generative scoring is a computationally efficient and superior alternative to standard contrastive scores like CLIPScore for evaluating image-text alignment. Crucially, they show that benchmark performance metrics may reflect linguistic artifacts rather than true visual reasoning, presenting risks for teams selecting or evaluating multimodal systems based on uncalibrated benchmark leaderboards.
Decision-makers should adopt generative scoring mechanisms for downstream multimodal retrieval and text-to-image evaluation while applying debiasing when candidate text distributions differ from natural pre-training data. Benchmarking teams must audit evaluation datasets to eliminate unnatural negative captions that allow blind models to succeed. Future development should explore advanced sampling techniques, address generative model training biases on long-tail data, and validate debiased generative scoring across larger multimodal architectures.
- Paper: Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality, Anuj Diwan et al. (2022). Its analysis of Winoground’s compositionality failures provides essential context for the source’s use of that benchmark to test image–text matching beyond word-level cues.
- Paper: VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena, Letitia Parcalabescu et al. (2022). VALSE establishes diagnostic tests for whether models ground linguistic phenomena in images, framing the source’s investigation of when benchmark scores reflect language shortcuts instead.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). Its balanced VQA benchmark demonstrates how language-only shortcuts can inflate multimodal scores, a central concern in the source’s audit of benchmark linguistic bias.
No sufficiently relevant recommendations were found.
