When and why vision-language models behave like bags-of-words, and what to do about it?

Mert YuksekgonulFederico BianchiPratyusha KalluriDan JurafskyJames Y. Zou

article2022ICLR726 citationsOral (notable-top-5%)

Demonstrates why vision-language models fail to process word order and object relations, introducing the large-scale ARO benchmark to diagnose these compositional blind spots and a hard-negative training strategy to fix them.

Listen

Modern artificial intelligence systems increasingly combine computer vision and natural language to understand images, perform search, and generate new visual content. While leading vision-language models achieve impressive scores on standard performance benchmarks, questions remain about whether they genuinely understand how objects, attributes, and actions connect. When evaluating complex real-world scenes, distinguishing between phrases such as "the horse is eating the grass" and "the grass is eating the horse" is critical for reliable performance, safety, and fairness in practical applications.

The article introduces a large-scale evaluation framework called the Attribution, Relation, and Order (ARO) benchmark to systematically measure how well vision-language models understand compositional relationships, object attributes, and word order. The authors evaluate popular foundation models—including CLIP, BLIP, FLAVA, and X-VLM—and examine why current training and evaluation practices mask severe compositional weaknesses.

To conduct this evaluation, the researchers constructed ARO using over 50,000 test cases derived from established visual datasets. The benchmark tests three core abilities: linking descriptive properties to the correct objects, identifying directed spatial and action relationships between entities, and recognizing correct sentence order compared to systematically scrambled versions. In parallel, the authors performed controlled experiments on standard image-text retrieval benchmarks by measuring model performance when captions were word-shuffled or when images were sliced and rearranged into mismatched visual patches.

The investigation produced four primary findings. First, widely used foundation models frequently operate like simple "bags of words," often performing near or below chance level when tasked with identifying proper relational direction (for instance, BLIP selected the inverted caption "the grass is eating the horse" with 81% probability). Second, models demonstrated minimal sensitivity to grammatical structure and word order, frequently scoring near the 20% random guessing baseline on word-order tasks. Third, conventional cross-modal retrieval benchmarks were shown to be flawed measures of comprehension; models suffered only marginal performance losses on standard retrieval tasks even when word order or image patches were completely scrambled. Fourth, the authors found that incorporating "composition-aware hard negatives" during training—specifically pairing scenes with close visual neighbors and injecting captions with swapped linguistic elements—dramatically closed this performance gap. Fine-tuning CLIP with this technique improved relational understanding accuracy from 59% to 81% and caption-order accuracy from 46% to 86% without degrading general image classification or retrieval capabilities.

These findings indicate that existing models achieve high benchmark scores through shortcut learning rather than genuine linguistic or visual understanding because standard retrieval pretraining does not require models to resolve fine-grained compositional differences. This reliance on lexical shortcuts poses practical operational risks, especially in downstream systems such as text-to-image generators, which can produce unfaithful outputs or reinforce societal stereotypes when attribute bindings fail. Organizations relying on off-the-shelf vision-language models for precise visual search, automated surveillance, or generative tasks face unexpected errors if models merely detect keyword presence rather than actual context.

To address these deficiencies, practitioners and developers should immediately integrate hard-negative mining strategies into model training pipelines and adopt fine-grained compositional benchmarks like ARO alongside standard metrics. Relying solely on conventional retrieval accuracy should be discontinued when evaluating readiness for deployment in mission-critical environments. Further research should prioritize pretraining full-scale foundation models with composition-aware objectives from scratch, as current empirical tests were limited to targeted fine-tuning on a single compute node with smaller batch sizes. Overall, the benchmark results provide high confidence that mainstream vision-language models have severe compositional blind spots, which can be substantially corrected with modest, targeted adjustments to contrastive training data.

Cover for When and why vision-language models behave like bags-of-words, and what to do about it?

Abstract

Despite the success of large vision and language models (VLMs) in many downstream applications, it is unclear how well they encode compositional information. Here, we create the Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the ability of VLMs to understand different types of relationships, attributes, and order. ARO consists of Visual Genome Attribution, to test the understanding of objects' properties; Visual Genome Relation, to test for relational understanding; and COCO & Flickr30k-Order, to test for order sensitivity. ARO is orders of magnitude larger than previous benchmarks of compositionality, with more than 50,000 test cases. We show where state-of-the-art VLMs have poor relational understanding, can blunder when linking objects to their attributes, and demonstrate a severe lack of order sensitivity. VLMs are predominantly trained and evaluated on large datasets with rich compositional structure in the images and captions. Yet, training on these datasets has not been enough to address the lack of compositional understanding, and evaluating on these datasets has failed to surface this deficiency. To understand why these limitations emerge and are not represented in the standard tests, we zoom into the evaluation and training procedures. We demonstrate that it is possible to perform well on retrieval over existing datasets without using the composition and order information. Given that contrastive pretraining optimizes for retrieval on datasets with similar shortcuts, we hypothesize that this can explain why the models do not need to learn to represent compositional information. This finding suggests a natural solution: composition-aware hard negative mining. We show that a simple-to-implement modification of contrastive learning significantly improves the performance on tasks requiring understanding of order and compositionality.

Table of Contents

  • 1 Introduction
  • 2 Attribution, Relation, and Order (ARO) Benchmark: When do models behave like a bag-of-words?
  • 2.1 New Benchmarks for Assessing Relational and Attributive Understanding
  • 2.2 New Benchmarks for Assessing Order Sensitivity
  • 2.3 Evaluating VLMs on ARO
  • 3 Why do models behave like bag-of-words? A critique of retrieval and contrastive pretraining
  • 3.1 Limitations of retrieval as an evaluation
  • 3.2 Limitations of retrieval and contrastive pretraining as an objective
  • 4 A simple fix: Composition-aware hard negatives
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Proposed Datasets
  • A.1 Generating the Datasets
  • B Models
  • C Negative Mining
  • C.1 Negative Text Mining
  • C.2 Negative Image Mining
  • C.3 Fine-tuning Details

Knowls

  1. Knowl 1 — ARO tests object relations, attributes, and caption order at scale

    definition

    The Attribution, Relation, and Order (ARO) benchmark evaluates whether vision-language models distinguish compositional differences in images and captions. Its Visual Genome Relation task presents a cropped image with a correct caption of the form “the [object 1] is [relation] [object 2]” and a caption with the two objects swapped. Its Visual Genome Attribution task tests whether a model assigns two different attributes to the right objects, contrasting captions such as “the [attribute 1] [object 1] and the [attribute 2] [object 2]” with captions that swap the attributes. These Visual Genome tasks contain 23,937 relation cases spanning 48 relations and 28,748 attribution cases spanning 117 attribute pairs.

    The Visual Genome cases are generated from GQA scene graphs. Candidate objects must each be at least one quarter of the image’s width and height; relation pairs must be from different object categories, and attribution pairs must have distinct objects and distinct attributes. The benchmark crops the smallest box containing both objects, uses the templates above to form the correct and swapped captions, and removes symmetric relations from the relation task.

    ARO also includes COCO-Order and Flickr30k-Order. Each case pairs an image with its original caption and four perturbed versions: nouns and adjectives are shuffled, everything except nouns and adjectives is shuffled, trigrams are shuffled, or the words within each trigram are shuffled. SpaCy part-of-speech tagging is used to support the perturbations. The Visual Genome tasks together provide 52,685 cases; the order tasks test selection among five captions, for which chance accuracy is 20%.

  2. Knowl 2 — Relational understanding varies by model and relation type

    empirical result

    On Visual Genome Relation, macro accuracy was 0.59 for CLIP, 0.80 for NegCLIP, 0.64 for CLIP fine-tuned on COCO without the proposed hard negatives (CLIP-FT), 0.73 for X-VLM, 0.59 for BLIP, and 0.24 for Flava. Chance accuracy is 0.50 because each image is paired with a correct and a swapped caption. The category-level accuracies show that performance is not uniform across relation types: spatial-relation accuracy was 0.56 for CLIP, 0.66 for NegCLIP, 0.57 for CLIP-FT, 0.74 for X-VLM, 0.66 for BLIP, and 0.34 for Flava; verb-relation accuracy was 0.61, 0.86, 0.66, 0.73, 0.56, and 0.20, respectively. Thus, several models are near or below chance overall, and Flava is below chance in both categories.

  3. Knowl 3 — Attribution accuracy differs sharply across vision-language models

    empirical result

    On Visual Genome Attribution, macro accuracy was 0.62 for CLIP, 0.65 for CLIP-FT, 0.71 for NegCLIP, 0.87 for X-VLM, 0.88 for BLIP, and 0.73 for Flava. Each case asks a model to choose between the correct assignment of two attributes to two objects and a caption with those assignments swapped; chance accuracy is 0.50. The results therefore show a substantial model-to-model range: CLIP is relatively close to chance, while BLIP and X-VLM perform much better on this task.

  4. Knowl 4 — Models show weak and uneven preference for correctly ordered captions

    empirical result

    In the five-choice Pick the Right Caption evaluation, chance accuracy is 0.20. On Flickr30k-PRC, accuracy was 0.369 ± 0.009 for BLIP, 0.595 ± 0.006 for CLIP, 0.129 ± 0.005 for Flava, and 0.473 ± 0.003 for X-VLM. On COCO-PRC, it was 0.321 ± 0.001 for BLIP, 0.460 ± 0.001 for CLIP, 0.039 ± 0.001 for Flava, and 0.362 ± 0.003 for X-VLM. The task compares an image’s original caption against four systematic word-order perturbations. These results show that performance is model- and dataset-dependent; Flava is below chance on both datasets, while the other models range from modest to stronger-than-chance performance.

  5. Knowl 5 — Retrieval remains possible after caption or image order is disrupted

    empirical result

    The authors tested whether standard cross-modal retrieval requires compositional order by perturbing captions or image layouts in the COCO and Flickr30k Karpathy test splits, containing 5,000 and 1,000 images, respectively. Retrieval was measured with text and image Recall@1 and Recall@5. Caption perturbations included shuffling all words and other systematic permutations; image perturbations divided images into four rows, four columns, or nine patches and shuffled the resulting pieces.

    For example, after shuffling all caption words in COCO, text Recall@1 was 0.690 ± 0.004 for BLIP, 0.341 ± 0.001 for CLIP, 0.338 ± 0.006 for Flava, and 0.633 ± 0.006 for X-VLM. On Flickr30k, the corresponding scores were 0.902 ± 0.008, 0.587 ± 0.007, 0.504 ± 0.002, and 0.869 ± 0.003. Even severe image perturbations often left retrieval well above zero: after shuffling nine patches on COCO, BLIP’s text Recall@1 was 0.594 ± 0.003 and its image Recall@1 was 0.488 ± 0.002; on Flickr30k, these were 0.875 ± 0.004 and 0.739 ± 0.001. Across the tested perturbations and models, retrieval performance generally declined but remained substantial, showing that strong retrieval scores do not by themselves establish sensitivity to composition or order.

  6. Knowl 6 — Retrieval training can reward a bag-of-words shortcut

    assumption

    The paper argues that contrastive pretraining may not incentivize models to encode word order or compositional structure when the training data offers few hard alternatives. Contrastive image-text training optimizes matching between paired images and captions, much like retrieval. The retrieval perturbation experiments show that high retrieval performance can persist when order cues are damaged. The authors therefore hypothesize that, in datasets designed to cover broad semantic concepts but containing few examples with similar words that must be distinguished by their ordering, a model can obtain high reward without representing composition. Under those conditions, behaving like a bag of words is a viable shortcut, rather than a demonstrated necessity of contrastive learning.

  7. Knowl 7 — NegCLIP adds composition-aware text and image negatives

    model/method

    NegCLIP modifies CLIP-style contrastive training to make fine-grained differences harder to ignore. For each training caption, the method uses spaCy to generate candidate negative captions by swapping elements such as nouns, adjectives, adverbs, verb phrases, or noun phrases; noun phrases are used only when they contain at least three tokens, avoiding overlap with noun swaps. Captions with no available negative are removed. At each epoch, one available negative caption is sampled for each caption and added to the training batch.

    The method also mines image negatives from the COCO training images. It computes pairwise CLIP image similarities, stores the three nearest neighbors for each image, and samples one of those neighbors per image at each epoch; the sampled image and its associated captions and negative captions are added to the batch. In the contrastive objective, the original captions and negative captions are concatenated, producing image-to-caption similarities for both sets. The row-wise and column-wise cross-entropy losses are computed as in CLIP, except that no column-wise loss is applied to negative captions, which have no matching image. The targeted caption negatives challenge order and composition, while similar-image negatives challenge discrimination between visually related scenes.

  8. Knowl 8 — NegCLIP substantially improves compositional evaluations with limited downstream losses

    empirical result

    After fine-tuning, NegCLIP improved all four reported compositional evaluation scores over the original CLIP: Visual Genome Relation rose from 0.59 to 0.81, Visual Genome Attribution from 0.62 to 0.71, Flickr30k-PRC from 0.59 to 0.91, and COCO-PRC from 0.46 to 0.86. The COCO fine-tuning control without hard negatives scored 0.63, 0.65, 0.50, and 0.36 on those tasks, respectively, indicating that ordinary fine-tuning alone did not account for the full gains.

    The reported downstream metrics were CIFAR10, CIFAR100, and ImageNet accuracy, plus Recall@1 for image-to-text and text-to-image retrieval on Flickr30k and COCO. CLIP versus NegCLIP scores were: CIFAR10, 0.95 versus 0.94; CIFAR100, 0.80 versus 0.79; ImageNet, 0.75 versus 0.72; Flickr30k image Recall@1, 0.59 versus 0.67; Flickr30k text Recall@1, 0.78 versus 0.79; COCO image Recall@1, 0.30 versus 0.41; and COCO text Recall@1, 0.50 versus 0.56. Thus, the reported compositional gains came without substantial losses on these downstream measurements.

  9. Knowl 9 — NegCLIP was evaluated as a resource-constrained COCO fine-tuning study

    experimental setup

    Because training CLIP from scratch was computationally infeasible for the study, the authors fine-tuned the ViT-B/32 CLIP model on the COCO training split. They trained both NegCLIP and a CLIP-FT control without sampled hard negatives for five epochs, using AdamW, 50 warmup steps, a cosine-annealing learning-rate schedule, batch size 1,024, and one NVIDIA RTX 2080 Ti GPU. They swept learning rates of 1e-5, 5e-6, and 1e-6 and selected models using retrieval performance on the COCO validation split. Evaluation included the four compositional tasks and downstream classification and retrieval tasks.

  10. Knowl 10 — The hard-negative result is limited to fine-tuning and a smaller batch regime

    limitation

    The paper did not train a vision-language model from scratch with composition-aware hard negatives, so its experiments establish improvements from fine-tuning rather than from composition-aware contrastive pretraining. The authors identify verification during pretraining as future work. Their batch size was 1,024 on a single GPU, compared with the batch size of 32,000 used by the cited CLIP training, and they expect that larger batch sizes could yield further improvements.

Coverage note — No substantial contributed material is omitted; background and related work are excluded, while the principal benchmark, retrieval analysis, hard-negative method, results, and stated training limitations are represented.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022. 1
  2. 2.Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions? ArXiv preprint, abs/2210.15230, 2022. URL https://arxiv.org/abs/2210.15230. 10
  3. 3.Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale, 2022. URL https://arxiv.org/abs/2211.03759. 10
  4. 4.Abeba Birhane and Vinay Uday Prabhu. Large image datasets: A pyrrhic win for computer vision? In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1536–1546. IEEE, 2021. 10
  5. 5.Ben Bogin, Shivanshu Gupta, Matt Gardner, and Jonathan Berant. Covr: A test-bed for visually grounded compositional generalization with real images. arXiv preprint arXiv:2109.10613, 2021. 8
  6. 6.Wieland Brendel and Matthias Bethge. Approximating CNNs with bag-of-local-features models works surprisingly well on imagenet. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SkfMWhAqYQ. 9
  7. 7.Meredith Broussard. Artificial unintelligence: How computers misunderstand the world. mit Press, 2018. 10
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020. 6
  9. 9.Jaemin Cho, Abhaysinh Zala, and Mohit Bansal. DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generative Transformers. ArXiv preprint, abs/2202.04053, 2022. URL https://arxiv.org/abs/2202.04053. 10
  10. 10.Colin Conwell and Tomer D Ullman. Testing relational understanding in text-guided image generation. arXiv preprint arXiv:2208.00005, 2022. 5, 8
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009. 7
  12. 12.Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald. Why is winoground hard? investigating failures in visuolinguistic compositionality. arXiv preprint arXiv:2211.00768, 2022. 2, 8
  13. 13.Allyson Ettinger. What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models. Transactions of the Association for Computational Linguistics, 8:34–48, 2020. 9
  14. 14.Stella Frank, Emanuele Bugliarello, and Desmond Elliott. Vision-and-language or vision-for-language? On cross-modal influence in multimodal transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021). Association for Computational Linguistics, nov 2021. URL https://arxiv.org/abs/2109.04448. 8
  15. 15.Weifeng Ge. Deep metric learning with hierarchical triplet loss. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 269–285, 2018. 9
  16. 16.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bygh9j09KX. 7
  17. 17.Robert Geirhos, Jorn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020. 2, 6
  18. 18.Ben Harwood, Vijay Kumar BG, Gustavo Carneiro, Ian Reid, and Tom Drummond. Smart mining for deep metric learning. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2821–2829, 2017. 9
  19. 19.Jack Hessel and Alexandra Schofield. How effective is bert without word ordering? implications for language understanding and data privacy. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pp. 204–211, 2021. 9
  20. 20.Matthew Honnibal and Ines Montani. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear, 2017. 4
  21. 21.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6700–6709, 2019. 3, 8, 16
  22. 22.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pp. 4904–4916. PMLR, 2021. 2
  23. 23.Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus. Hard negative mixing for contrastive learning. Advances in Neural Information Processing Systems, 33:21798–21809, 2020. 7, 9
  24. 24.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3128–3137, 2015. 5
  25. 25.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73, 2017. 3, 8, 10
  26. 26.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 7
  27. 27.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34:9694–9705, 2021. 6, 9
  28. 28.Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022. 1, 4, 5, 6, 9, 18
  29. 29.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014. 4, 5
  30. 30.Joe O’Connor and Jacob Andreas. What context features can transformer language models use? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 851–864, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.70. URL https://aclanthology.org/2021.acl-long.70. 4, 9
  31. 31.OpenAI. DALL·E Now Available Without Waitlist — openai.com. https://openai.com/blog/dall-e-now-available-without-waitlist/, 2022. [Accessed 01-Nov-2022]. 10
  32. 32.Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, and Albert Gatt. Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena. arXiv preprint arXiv:2112.07566, 2021a. 8
  33. 33.Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks. In Proceedings of the 1st Workshop on Multimodal Semantic Representations (MMSR), pp. 32–44, Groningen, Netherlands (Online), June 2021b. Association for Computational Linguistics. URL https://aclanthology.org/2021.mmsr-1.4. 8
  34. 34.Kenny Peng, Arunesh Mathur, and Arvind Narayanan. Mitigating dataset harms requires stewardship: Lessons from 1000 papers. arXiv preprint arXiv:2108.02922, 2021. 10
  35. 35.Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 1145–1160, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.98. URL https://aclanthology.org/2021.findings-acl.98. 9
  36. 36.Yao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan, Alex Beutel, and Xuezhi Wang. Understanding and improving robustness of vision transformers through patch-based negative augmentation. arXiv preprint arXiv:2110.07858, 2021. 6, 9
  37. 37.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021. 1, 2, 4, 6, 17, 20
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020. 5
  39. 39.Inioluwa Deborah Raji and Genevieve Fried. About face: A survey of facial recognition evaluation. arXiv preprint arXiv:2102.00813, 2021. 10
  40. 40.Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. Contrastive learning with hard negative samples. ArXiv, abs/2010.04592, 2021. 7, 9
  41. 41.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. 5, 9
  42. 42.Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15638–15650, 2022. 1, 4, 19
  43. 43.Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. In EMNLP, 2021. 9
  44. 44.Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. A corpus of natural language for visual reasoning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 217–223, 2017. 8
  45. 45.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. arXiv preprint arXiv:1811.00491, 2018. 8
  46. 46.Ajinkya Tejankar, Bichen Wu, Saining Xie, Madian Khabsa, Hamed Pirsiavash, and Hamed Firooz. A fistful of words: Learning transferable visual models from bag-of-words supervision. arXiv preprint arXiv:2112.13884, 2021. 9
  47. 47.Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5238–5248, 2022. 1, 2, 8
  48. 48.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018. 9
  49. 49.Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan. Omnivl: One foundation model for image-language and video-language tasks. arXiv preprint arXiv:2209.07526, 2022a. 1
  50. 50.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022b. 1
  51. 51.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6. 19
  52. 52.Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. Contrastive language-vision ai models pretrained on web-scraped multimodal data exhibit sexual objectification bias. arXiv preprint arXiv:2212.11261, 2022. 10
  53. 53.Chao-Yuan Wu, R Manmatha, Alexander J Smola, and Philipp Krahenbuhl. Sampling matters in deep embedding learning. In Proceedings of the IEEE international conference on computer vision, pp. 2840–2848, 2017. 9
  54. 54.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014. 4, 5
  55. 55.Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. ArXiv, abs/2111.08276, 2022a. 6
  56. 56.Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 25994–26009. PMLR, 2022b. 4, 18
  57. 57.Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre. A benchmark for compositional visual reasoning. arXiv preprint arXiv:2206.05379, 2022. 8
  58. 58.Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18123–18133, 2022. 1
  59. 59.Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. Contrastive learning of medical visual representations from paired images and text. arXiv preprint arXiv:2010.00747, 2020. 2, 6

Citation

MLA
Yuksekgonul, M., et al. “When and Why Vision-language Models Behave Like Bags-of-words, and What to Do About It?”. arXiv, 2022, http://arxiv.org/abs/2210.01936v3.
APA
Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., & Zou, J. (2022). When and why vision-language models behave like bags-of-words, and what to do about it?. arXiv. http://arxiv.org/abs/2210.01936v3
Chicago
Yuksekgonul, M., F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou. 2022. “When and Why Vision-language Models Behave Like Bags-of-words, and What to Do About It?”. arXiv. http://arxiv.org/abs/2210.01936v3.
Harvard
Yuksekgonul, M. et al. (2022) “When and why vision-language models behave like bags-of-words, and what to do about it?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.01936v3.
Vancouver
1. Yuksekgonul M, Bianchi F, Kalluri P, Jurafsky D, Zou J (2022) When and why vision-language models behave like bags-of-words, and what to do about it?. arXiv

BibTeX

@article{yuksekgonul2022when,
  title = {When and why vision-language models behave like bags-of-words, and what to do about it?},
  author = {Yuksekgonul, Mert and Bianchi, Federico and Kalluri, Pratyusha and Jurafsky, Dan and Zou, James},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.01936v3},
  eprint = {2210.01936}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/