Revisiting the Role of Language Priors in Vision-Language Models

Zhiqiu LinXinyue ChenDeepak PathakPengchuan ZhangDeva Ramanan

article2024ICML56 citations

Introduces VisualGPTScore to repurpose generative vision-language models for zero-shot retrieval tasks and provides a training-free debiasing method that corrects for linguistic priors to achieve state-of-the-art accuracy across vision-language benchmarks.

Listen

Vision-language models are widely adopted for multimodal understanding tasks because they can operate out-of-the-box without task-specific fine-tuning. However, standard contrastive models often struggle with complex compositional reasoning—such as distinguishing "the horse is eating the grass" from "the grass is eating the horse"—while existing evaluation benchmarks frequently suffer from hidden linguistic biases that obscure true visual understanding.

The article aims to evaluate whether generative vision-language models can be repurposed for matching tasks and demonstrates a probabilistic method to control for language bias without model retraining.

The authors analyze image-conditioned generative models across nine retrieval and alignment benchmarks, including ARO, SugarCrepe, VL-CheckList, COCO, Flickr30K, and Winoground. They introduce the Visual Generative Pre-Training Score, which uses the model's conditional probability of generating a text caption given an image as a retrieval matching score. To resolve mismatches between pre-training language distributions and test distributions, they introduce a training-free debiasing formula that adjusts match scores using an estimated language prior derived via efficient Monte Carlo sampling of Gaussian noise images.

The analysis reveals four central findings. First, off-the-shelf generative scoring consistently outperforms traditional discriminative models and heavily engineered baselines across compositionality benchmarks without requiring additional training data. Second, multiple widely used benchmarks contain unnatural negative captions that can be solved by "blind" language-only models without looking at images, achieving up to 98–99% accuracy on benchmarks like ARO. Third, on realistic benchmarks where negative captions are plausible, raw generative models can wrongly favor common phrases; applying the proposed debiasing technique resolves this bias and improves accuracy, for example boosting Winoground text scores from 27.5% to 36.6% and ImageNet zero-shot classification from 18.6% to 40.0%. Fourth, applying debiased generative scoring to modern vision-language models like LLaVA-1.5 sets a new state-of-the-art on demanding image-text alignment tasks, outperforming complex pipelines that rely on auxiliary models like ChatGPT.

These findings indicate that generative scoring is a computationally efficient and superior alternative to standard contrastive scores like CLIPScore for evaluating image-text alignment. Crucially, they show that benchmark performance metrics may reflect linguistic artifacts rather than true visual reasoning, presenting risks for teams selecting or evaluating multimodal systems based on uncalibrated benchmark leaderboards.

Decision-makers should adopt generative scoring mechanisms for downstream multimodal retrieval and text-to-image evaluation while applying debiasing when candidate text distributions differ from natural pre-training data. Benchmarking teams must audit evaluation datasets to eliminate unnatural negative captions that allow blind models to succeed. Future development should explore advanced sampling techniques, address generative model training biases on long-tail data, and validate debiased generative scoring across larger multimodal architectures.

arXiv: 2306.01879

No sufficiently relevant recommendations were found.

Cover for Revisiting the Role of Language Priors in Vision-Language Models

Abstract

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study generative VLMs that are trained for next-word generation given an image. We explore their zero-shot performance on the illustrative task of image-text retrieval across nine popular vision-language benchmarks. Our first observation is that they can be repurposed for discriminative tasks (such as image-text retrieval) by simply computing the match score of generating a particular text string given an image. We call this probabilistic score the Visual Generative Pre-Training Score (VisualGPTScore). While the VisualGPTScore produces near-perfect accuracy on some retrieval benchmarks, it yields poor accuracy on others. We analyze this behavior through a probabilistic lens, pointing out that some benchmarks inadvertently capture unnatural language distributions by creating adversarial but unlikely text captions. In fact, we demonstrate that even a “blind” language model that ignores any image evidence can sometimes outperform all prior art, reminiscent of similar challenges faced by the visual-question answering (VQA) community many years ago. We derive a probabilistic post-processing scheme that controls for the amount of linguistic bias in generative VLMs at test time without having to retrain or fine-tune the model. We show that the VisualGPTScore, when appropriately debiased, is a strong zero-shot baseline for vision-language understanding, oftentimes producing state-of-the-art accuracy.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 The role of language priors
  • 4 Experimental results on I-to-T retrieval
  • 5 Additional Challenging Benchmarks
  • 6 Discussion and Limitations
  • References
  • A Is VisualGPTScore a Biased Estimator of Pt​r​a​i​n​(𝐭|𝐢)P_{train}(\mathbf{t}|\mathbf{i})?
  • B Ablation Studies on α\alpha-Debiasing
  • C Experiments with BLIP-2
  • D Additional Reports
  • E Benchmark Visualization

Knowls

  1. Knowl 1 — VisualGPTScore turns image-conditioned text generation into an image–text matching score

    model/method

    For an image ii and a candidate caption t=(t_1,ots,t_m), an image-conditioned autoregressive language model assigns the caption likelihood P(t∣i)=∏k=1mP(tk∣t<k,i)P(t\mid i)=\prod_{k=1}^{m}P(t_k\mid t_{<k},i). The paper uses its length-normalized geometric mean as the Visual Generative Pre-Training Score (VisualGPTScore):

    VisualGPTScore⁡(t,i)=exp⁡ ⁣(1m∑k=1mlog⁡P(tk∣t<k,i)).\operatorname{VisualGPTScore}(t,i)=\exp\!\left(\frac{1}{m}\sum_{k=1}^{m}\log P(t_k\mid t_{<k},i)\right).

    Here mm is the number of caption tokens and t<kt_{<k} denotes the tokens preceding position kk. Given the full candidate caption, all next-token probabilities can be evaluated in parallel rather than generating tokens sequentially. The score repurposes the model’s generation head for retrieval without fine-tuning; in the BLIP implementation it has the same computational cost as BLIP’s image–text matching score.

  2. Knowl 2 — Bayesian language-prior reweighting corrects train–test caption shift

    model/method

    Let Ptrain(t,i)P_{\mathrm{train}}(t,i) and Ptest(t,i)P_{\mathrm{test}}(t,i) be the training and test joint distributions over captions tt and images ii. The paper assumes that the image likelihood is unchanged, Ptest(i∣t)=Ptrain(i∣t)P_{\mathrm{test}}(i\mid t)=P_{\mathrm{train}}(i\mid t), while the caption prior may shift. Under this assumption, the test-optimal image-to-text ranking score is

    Ptest(t∣i)∝Ptrain(t∣i) Ptest(t)Ptrain(t).P_{\mathrm{test}}(t\mid i)\propto P_{\mathrm{train}}(t\mid i)\,\frac{P_{\mathrm{test}}(t)}{P_{\mathrm{train}}(t)}.

    Because the test caption prior is often unavailable, the paper uses a one-parameter approximation: Ptest(t)∝Ptrain(t)1−αP_{\mathrm{test}}(t)\propto P_{\mathrm{train}}(t)^{1-\alpha}, with 0≤α≤10\leq\alpha\leq1. This gives the ranking score Ptrain(t∣i)/Ptrain(t)αP_{\mathrm{train}}(t\mid i)/P_{\mathrm{train}}(t)^\alpha. At α=0\alpha=0 the raw VisualGPTScore is retained, appropriate when train and test caption priors match; at α=1\alpha=1 the training caption prior is fully removed, corresponding to an uninformative test prior. A held-out validation set can be used to choose an intermediate α\alpha.

  3. Knowl 3 — A small number of Gaussian-noise images estimates the training caption prior

    algorithm

    To estimate the marginal caption probability needed for prior reweighting, the paper uses the identity Ptrain(t)=Ei∼Ptrain(i)[Ptrain(t∣i)]P_{\mathrm{train}}(t)=\mathbb{E}_{i\sim P_{\mathrm{train}}(i)}[P_{\mathrm{train}}(t\mid i)]. With nn sampled images i1,…,ini_1,\ldots,i_n, the Monte Carlo estimate is P^train(t)=1n∑k=1nPtrain(t∣ik)\widehat P_{\mathrm{train}}(t)=\frac{1}{n}\sum_{k=1}^{n}P_{\mathrm{train}}(t\mid i_k). Instead of drawing images from the training corpus, the proposed efficient procedure evaluates the caption likelihood on “null” Gaussian-noise images. For BLIP and BLIP-2, the reported noise has mean 1.01.0 and standard deviation 0.250.25; the default sample count is 3 for most benchmarks, 100 for Winoground, 30 for EqBen, and 1 for ImageNet. The paper reports that as few as 1–3 such images can give robust estimates and that null-image estimates can be more sample-efficient and less variable than estimates from training images.

  4. Knowl 4 — Raw and debiased VisualGPTScore perform strongly on four compositional retrieval suites

    empirical result

    Using the off-the-shelf BLIP model with a ViT-L image encoder, pretrained on LAION-114M and not fine-tuned for these tests, the paper reports image-to-text retrieval accuracy on ARO, VL-CheckList, SugarCrepe, and Crepe. In each sequence below, the values are ordered as named in the benchmark description; the three reported scores are raw VisualGPTScore (α=0\alpha=0), full debiasing (α=1\alpha=1), and validation-selected α∗\alpha^*, respectively.

    • ARO (VG-Relation, VG-Attribution, COCO-Order, Flickr30K-Order): raw 89.1, 95.3, 99.4, 99.5; α=1\alpha=1: 68.1, 87.9, 32.4, 44.5; α∗\alpha^*: 89.1, 95.4, 99.4, 99.5.
    • VL-CheckList (Object, Attribute, Relation): raw 92.6, 78.7, 90.8; α=1\alpha=1: 90.4, 77.6, 77.8; α∗\alpha^*: 94.4, 82.1, 92.8.
    • SugarCrepe (Replace, Swap, Add): raw 93.3, 91.0, 91.0; α=1\alpha=1: 83.2, 85.5, 85.9; α∗\alpha^*: 95.1, 92.4, 97.4.
    • Crepe (Atom, Swap, Negate): raw 73.2, 78.1, 79.6; α=1\alpha=1: 20.6, 28.3, 35.6; α∗\alpha^*: 73.3, 78.1, 79.6.

    All values are percentages. The results show that raw generation scores can already be very strong, while the best amount of prior removal depends on benchmark construction: full debiasing can damage performance on some suites, whereas validation-selected debiasing improves SugarCrepe and VL-CheckList.

  5. Knowl 5 — Image-blind language scores expose caption-distribution artifacts in retrieval benchmarks

    empirical result

    The paper evaluates captions using a BLIP-derived estimate of Ptrain(t)P_{\mathrm{train}}(t), obtained with Gaussian-noise images so that the score contains no relevant image evidence. This blind score achieves 87.6, 80.7, 98.6, and 99.1 percent on ARO’s VG-Relation, VG-Attribution, COCO-Order, and Flickr30K-Order tasks, respectively; random chance is 50, 50, 20, and 20 percent. It also reaches 75.9, 77.1, and 70.9 percent on SugarCrepe Replace, Swap, and Add, against 50 percent chance, and 55.4, 69.7, and 60.8 percent on Crepe Atom, Swap, and Negate, against 16.7 percent chance.

    These results demonstrate that some retrieval benchmarks can be substantially solved from caption plausibility or frequency alone. This is especially pronounced when negative captions are implausible, such as word-shuffled captions in ARO; realistic, carefully curated negatives reduce this advantage but do not eliminate it in every benchmark.

  6. Knowl 6 — Fixed prior removal improves VisualGPTScore on five additional challenging tasks

    empirical result

    On Winoground, EqBen, COCO retrieval, Flickr30K retrieval, and ImageNet-1K, raw BLIP VisualGPTScore is weak, but fixed α=1\alpha=1 debiasing improves all five results. The reported raw, fixed-α=1\alpha=1, and validation-selected-α\alpha results are:

    • Winoground text score: 27.5 (2.3), 33.7 (2.4), and 36.6 (2.6); selected α=0.855\alpha=0.855 (0.023).
    • EqBen text score: 9.6 (0.2), 19.8 (0.3), and 19.8 (0.3); selected α=0.992\alpha=0.992 (0.007).
    • COCO I-to-T Recall@1/Recall@5: 19.7/40.6, 46.2/73.1, and 48.0/74.2; selected α=0.819\alpha=0.819.
    • Flickr30K I-to-T Recall@1/Recall@5: 34.6/59.0, 58.7/88.0, and 63.6/89.2; selected α=0.719\alpha=0.719.
    • ImageNet-1K zero-shot classification accuracy: 18.6, 36.2, and 40.0 percent; selected α=0.670\alpha=0.670.

    The Winoground and EqBen results are mean scores with standard deviations in parentheses. For those datasets, the authors selected α\alpha by grid search on half the data and evaluated on the other half, repeating the split ten times; COCO and Flickr30K used their official validation sets, and ImageNet used one-shot validation samples. The gains show that removing a language prior can help when incorrect captions are plausible or caption frequencies are not a useful proxy for image match.

  7. Knowl 7 — Language-prior debiasing is related to a tunable pointwise mutual-information score

    theoretical result

    For image-to-text retrieval, the paper identifies its α\alpha-debiased ranking with a PMIk-style association score. Let P(t,i)P(t,i) be the joint probability of caption tt and image ii, and let P(t)P(t) and P(i)P(i) be their marginals. For α>0\alpha>0, the debiased score P(t∣i)/P(t)αP(t\mid i)/P(t)^\alpha induces the same candidate-caption ranking, up to factors constant for a fixed query image and a monotone transformation, as

    PMI⁡k(t,i)=P(t,i)kP(i)P(t),k=1α.\operatorname{PMI}_k(t,i)=\frac{P(t,i)^k}{P(i)P(t)},\qquad k=\frac{1}{\alpha}.

    Thus the amount of caption-prior correction corresponds to the strength parameter in this PMI-style score. The connection gives an information-theoretic interpretation of debiasing as suppressing the influence of marginal caption frequency; it also relates the method to an established way of controlling how strongly marginal probabilities affect association scores.

  8. Knowl 8 — The image-conditioned language model also yields a text-to-image retrieval score

    model/method

    For text-to-image retrieval, the Bayes-optimal test ranking is Ptest(i∣t)∝Ptrain(t∣i)Ptrain(i)P_{\mathrm{test}}(i\mid t)\propto P_{\mathrm{train}}(t\mid i)P_{\mathrm{train}}(i), where tt is a caption and ii is a candidate image. If the training image marginal Ptrain(i)P_{\mathrm{train}}(i) is approximately uniform over candidates, ranking images directly by the image-conditioned caption score Ptrain(t∣i)P_{\mathrm{train}}(t\mid i) is appropriate. With BLIP, the paper reports text-to-image results of 21.5 versus 15.8 for ITMScore on Winoground image score, and 26.1 versus 20.3 on EqBen image score. On COCO, VisualGPTScore achieves Recall@1/Recall@5 of 55.6/79.2 versus ITMScore’s 54.8/79.0; on Flickr30K it reaches 76.8/93.4 versus 77.8/93.9. The results are competitive with ITMScore, though the paper notes that text-to-image retrieval is less affected by caption-prior bias.

  9. Knowl 9 — LLaVA-1.5 VisualGPTScore improves image–text alignment evaluations

    empirical result

    On Winoground and EqBen-Mini, the paper applies VisualGPTScore to the off-the-shelf LLaVA-1.5-13B model. Scores are reported as text/image/group accuracy percentages. Raw VisualGPTScore (α=0\alpha=0) obtains 36.3/37.0/24.8 on Winoground and 25.7/42.1/21.4 on EqBen-Mini. With α=1\alpha=1, the respective scores rise to 44.3/37.0/27.5 and 42.9/42.1/29.3. Debiasing uses one Gaussian-noise image with mean 0 and standard deviation 0.25. The debiased method exceeds the listed alternative methods on Winoground text and group scores and on EqBen-Mini text and group scores, supporting the use of generative VLM scores as image–text alignment metrics without adding a separate language model or retraining.

  10. Knowl 10 — VisualGPTScore can be biased even on data from the model’s training distribution

    limitation

    The paper cautions that VisualGPTScore need not estimate the true Ptrain(t∣i)P_{\mathrm{train}}(t\mid i): the learned score may favor common captions. In a proposed bias model, the estimated conditional is P^train(t∣i)=Ptrain(t∣i)Ptrain(t)β\widehat P_{\mathrm{train}}(t\mid i)=P_{\mathrm{train}}(t\mid i)P_{\mathrm{train}}(t)^\beta, where the unknown β\beta represents that preference; averaging this estimate over training images gives P^train(t)=Ptrain(t)1+β\widehat P_{\mathrm{train}}(t)=P_{\mathrm{train}}(t)^{1+\beta}, up to normalization. This model explains why nonzero debiasing can help even when train and test distributions match.

    In image-to-text retrieval on random LAION training subsets, raw VisualGPTScore accuracy fell from 59.0 percent with 100 candidate captions to 25.1 percent with 5,000 candidates. Validation-selected debiasing raised those results to 95.0 and 54.1 percent, respectively. The analysis indicates a limitation of treating the model’s conditional and its sampled marginal as unbiased probabilities; the paper does not estimate the unknown bias parameter directly. It also notes that the underlying web-trained models inherit noise and imbalance from their pretraining data.

Coverage note — Detailed BLIP-2 variant results, per-subcategory Winoground breakdowns, and individual benchmark prior-frequency plots are omitted because they provide corroborating ablations rather than additional core methods or conclusions.

References

  1. 1.Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. Nocaps: Novel object captioning at scale. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 8948–8957, 2019.
  2. 2.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022.
  3. 3.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp. 2425–2433, 2015.
  4. 4.Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155, 2003.
  5. 5.Bertolini, L., Weeds, J., and Weir, D. Testing large language models on compositionality and inference with phrase-level adjective-noun entailment. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 4084–4100, 2022.
  6. 6.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. Improving image generation with better captions. https://cdn.openai.com/papers/dall-e-3.pdf, 2023.
  7. 7.Brendel, W. and Bethge, M. Approximating cnns with bag-of-local-features models works surprisingly well on imagenet. arXiv preprint arXiv:1904.00760, 2019.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  9. 9.Cascante-Bonilla, P., Shehada, K., Smith, J. S., Doveh, S., Kim, D., Panda, R., Varol, G., Oliva, A., Ordonez, V., Feris, R., et al. Going beyond nouns with vision & language models using synthetic data. arXiv preprint arXiv:2303.17590, 2023.
  10. 10.Cho, J., Hu, Y., Garg, R., Anderson, P., Krishna, R., Baldridge, J., Bansal, M., Pont-Tuset, J., and Wang, S. Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-image generation. arXiv preprint arXiv:2310.18235, 2023a.
  11. 11.Cho, J., Zala, A., and Bansal, M. Visual programming for text-to-image generation and evaluation. arXiv preprint arXiv:2305.15328, 2023b.
  12. 12.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A. W., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., hsin Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416, 2022.
  13. 13.Daille, B. Approche mixte pour l’extraction automatique de terminologie: statistiques lexicales et filtres linguistiques. PhD thesis, Ph. D. thesis, Universite Paris 7, 1994. ´
  14. 14.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  15. 15.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  16. 16.Diwan, A., Berry, L., Choi, E., Harwath, D., and Mahowald, K. Why is winoground hard? investigating failures in visuolinguistic compositionality. arXiv preprint arXiv:2211.00768, 2022.
  17. 17.Doveh, S., Arbelle, A., Harary, S., Panda, R., Herzig, R., Schwartz, E., Kim, D., Giryes, R., Feris, R., Ullman, S., et al. Teaching structured vision&language concepts to vision&language models. arXiv preprint arXiv:2211.11733, 2022.
  18. 18.Doveh, S., Arbelle, A., Harary, S., Alfassy, A., Herzig, R., Kim, D., Giryes, R., Feris, R., Panda, R., Ullman, S., et al. Dense and aligned captions (dac) promote compositional reasoning in vl models. arXiv preprint arXiv:2305.19595, 2023.
  19. 19.Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y. Eva: Exploring the limits of masked visual representation learning at scale. arXiv preprint arXiv:2211.07636, 2022.
  20. 20.Fu, J., Ng, S.-K., Jiang, Z., and Liu, P. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166, 2023.
  21. 21.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6904–6913, 2017.
  22. 22.Guo, C., Zhao, B., and Bai, Y. Deepcore: A comprehensive library for coreset selection in deep learning. In Database and Expert Systems Applications: 33rd International Conference, DEXA 2022, Vienna, Austria, August 22–24, 2022, Proceedings, Part I, pp. 181–195. Springer, 2022.
  23. 23.Henning, C. A. and Ewerth, R. Estimating the information gap between textual and visual representations. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval, pp. 14–22, 2017.
  24. 24.Herzig, R., Mendelson, A., Karlinsky, L., Arbelle, A., Feris, R., Darrell, T., and Globerson, A. Incorporating structured representations into pretrained vision & language models using scene graphs. arXiv preprint arXiv:2305.06343, 2023.
  25. 25.Hessel, J. and Schofield, A. How effective is bert without word ordering? implications for language understanding and data privacy. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pp. 204–211, 2021.
  26. 26.Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021.
  27. 27.Hsieh, C.-Y., Zhang, J., Ma, Z., Kembhavi, A., and Krishna, R. Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality. arXiv preprint arXiv:2306.14610, 2023.
  28. 28.Hu, Y., Liu, B., Kasai, J., Wang, Y., Ostendorf, M., Krishna, R., and Smith, N. A. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. arXiv preprint arXiv:2303.11897, 2023.
  29. 29.Huang, Y., Tang, J., Chen, Z., Zhang, R., Zhang, X., Chen, W., Zhao, Z., Lv, T., Hu, Z., and Zhang, W. Structure-clip: Enhance multi-modal language representations with structure knowledge. arXiv preprint arXiv:2305.06152, 2023.
  30. 30.Kamath, A., Hessel, J., and Chang, K.-W. Text encoders are performance bottlenecks in contrastive vision-language models. arXiv preprint arXiv:2305.14897, 2023.
  31. 31.Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y. Decoupling representation and classifier for long-tailed recognition. arXiv preprint arXiv:1910.09217, 2019.
  32. 32.Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858, 2019.
  33. 33.Li, B., Lin, Z., Pathak, D., Li, J., Fei, Y., Wu, K., Xia, X., Zhang, P., Neubig, G., and Ramanan, D. Evaluating and improving compositional text-to-visual generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024.
  34. 34.Li, J. and Jurafsky, D. Mutual information and diverse decoding improve neural machine translation, 2016.
  35. 35.Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 110–119, San Diego, California, June 2016. Association for Computational Linguistics. doi: 10.18653/v1/N16-1014. URL https://aclanthology.org/N16-1014.
  36. 36.Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., and Hoi, S. C. H. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34:9694–9705, 2021.
  37. 37.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888–12900. PMLR, 2022.
  38. 38.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  39. 39.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: ´ Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014.
  40. 40.Lin, Z., Yu, S., Kuang, Z., Pathak, D., and Ramana, D. Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models. arXiv preprint arXiv:2301.06267, 2023.
  41. 41.Lin, Z., Pathak, D., Li, B., Li, J., Xia, X., Neubig, G., Zhang, P., and Ramanan, D. Evaluating text-to-visual generation with image-to-text generation. arXiv preprint arXiv:2404.01291, 2024.
  42. 42.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023.
  43. 43.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  44. 44.Lu, Y., Yang, X., Li, X., Wang, X. E., and Wang, W. Y. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation. arXiv preprint arXiv:2305.11116, 2023.
  45. 45.Ma, Z., Hong, J., Gul, M. O., Gandhi, M., Gao, I., and Krishna, R. Crepe: Can vision-language foundation models reason compositionally? arXiv preprint arXiv:2212.07796, 2022.
  46. 46.Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54(6):1–35, 2021.
  47. 47.Miech, A., Alayrac, J.-B., Laptev, I., Sivic, J., and Zisserman, A. Thinking fast and slow: Efficient text-to-visual retrieval with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9826–9836, 2021.
  48. 48.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  49. 49.Papadimitriou, I., Futrell, R., and Mahowald, K. When classifying grammatical role, bert doesn’t care about word order... except when it matters. arXiv preprint arXiv:2203.06204, 2022.
  50. 50.Parashar, S., Lin, Z., Liu, T., Dong, X., Li, Y., Ramanan, D., Caverlee, J., and Kong, S. The neglected tails of vision-language models. arXiv preprint arXiv:2401.12425, 2024.
  51. 51.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. 2019.
  52. 52.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  53. 53.Role, F. and Nadif, M. Handling the impact of low frequency events on co-occurrence based measures of word similarity. In Proceedings of the international conference on Knowledge Discovery and Information Retrieval (KDIR-2011). Scitepress, pp. 218–223, 2011.
  54. 54.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  55. 55.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022.
  56. 56.Shapiro, A. Monte carlo sampling methods. Handbooks in operations research and management science, 10:353–425, 2003.
  57. 57.Shrivastava, A., Selvaraju, R. R., Naik, N., and Ordonez, V. Clip-lite: information efficient visual representation learning from textual annotations. arXiv preprint arXiv:2112.07133, 2021.
  58. 58.Singh, H., Zhang, P., Wang, Q., Wang, M., Xiong, W., Du, J., and Chen, Y. Coarse-to-fine contrastive learning in image-text-graph space for improved vision-language compositionality. arXiv preprint arXiv:2305.13812, 2023.
  59. 59.Sinha, K., Jia, R., Hupkes, D., Pineau, J., Williams, A., and Kiela, D. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. arXiv preprint arXiv:2104.06644, 2021.
  60. 60.Tejankar, A., Sanjabi, M., Wu, B., Xie, S., Khabsa, M., Pirsiavash, H., and Firooz, H. A fistful of words: Learning transferable visual models from bag-of-words supervision. arXiv preprint arXiv:2112.13884, 2021.
  61. 61.Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., and Ross, C. Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5238–5248, 2022.
  62. 62.Tschannen, M., Kumar, M., Steiner, A., Zhai, X., Houlsby, N., and Beyer, L. Image captioners are scalable vision learners too. arXiv preprint arXiv:2306.07915, 2023.
  63. 63.Wang, T., Lin, K., Li, L., Lin, C.-C., Yang, Z., Zhang, H., Liu, Z., and Wang, L. Equivariant similarity for vision-language foundation models. arXiv preprint arXiv:2303.14465, 2023.
  64. 64.Wang, Z., Feng, B., Narasimhan, K., and Russakovsky, O. Towards unique and informative captioning of images. In European Conference on Computer Vision (ECCV), 2020.
  65. 65.Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7959–7971, 2022.
  66. 66.Wu, X., Deng, Z., and Russakovsky, O. Multimodal dataset distillation for image-text retrieval. arXiv preprint arXiv:2308.07545, 2023.
  67. 67.Yao, T., Mei, T., and Ngo, C.-W. Co-reranking by mutual reinforcement for image search. In Proceedings of the ACM international conference on image and video retrieval, pp. 34–41, 2010.
  68. 68.Yarom, M., Bitton, Y., Changpinyo, S., Aharoni, R., Herzig, J., Lang, O., Ofek, E., and Szpektor, I. What you see is what you read? improving text-image alignment evaluation. arXiv preprint arXiv:2305.10400, 2023.
  69. 69.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  70. 70.Yuan, W., Neubig, G., and Liu, P. Bartscore: Evaluating generated text as text generation. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 27263–27277. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/e4d2b6e6fdeca3e60e0f1a62fee3d9dd-Paper.pdf.
  71. 71.Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J. When and why vision-language models behave like bag-of-words models, and what to do about it? arXiv preprint arXiv:2210.01936, 2022.
  72. 72.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  73. 73.Zhao, T., Wallace, E., Feng, S., Klein, D., and Singh, S. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, 2021.
  74. 74.Zhao, T., Zhang, T., Zhu, M., Shen, H., Lee, K., Lu, X., and Yin, J. Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations. arXiv preprint arXiv:2207.00221, 2022.

Citation

MLA
Lin, Z., et al. “Revisiting the Role of Language Priors in Vision-Language Models”. arXiv, 2023, http://arxiv.org/abs/2306.01879v4.
APA
Lin, Z., Chen, X., Pathak, D., Zhang, P., & Ramanan, D. (2023). Revisiting the Role of Language Priors in Vision-Language Models. arXiv. http://arxiv.org/abs/2306.01879v4
Chicago
Lin, Z., X. Chen, D. Pathak, P. Zhang, and D. Ramanan. 2023. “Revisiting the Role of Language Priors in Vision-Language Models”. arXiv. http://arxiv.org/abs/2306.01879v4.
Harvard
Lin, Z. et al. (2023) “Revisiting the Role of Language Priors in Vision-Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.01879v4.
Vancouver
1. Lin Z, Chen X, Pathak D, Zhang P, Ramanan D (2023) Revisiting the Role of Language Priors in Vision-Language Models. arXiv

BibTeX

@article{lin2023revisiting,
  title = {Revisiting the Role of Language Priors in Vision-Language Models},
  author = {Lin, Zhiqiu and Chen, Xinyue and Pathak, Deepak and Zhang, Pengchuan and Ramanan, Deva},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.01879v4},
  eprint = {2306.01879}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/