When and why vision-language models behave like bags-of-words, and what to do about it?
Mert YuksekgonulFederico BianchiPratyusha KalluriDan JurafskyJames Y. Zou
Demonstrates why vision-language models fail to process word order and object relations, introducing the large-scale ARO benchmark to diagnose these compositional blind spots and a hard-negative training strategy to fix them.
Modern artificial intelligence systems increasingly combine computer vision and natural language to understand images, perform search, and generate new visual content. While leading vision-language models achieve impressive scores on standard performance benchmarks, questions remain about whether they genuinely understand how objects, attributes, and actions connect. When evaluating complex real-world scenes, distinguishing between phrases such as "the horse is eating the grass" and "the grass is eating the horse" is critical for reliable performance, safety, and fairness in practical applications.
The article introduces a large-scale evaluation framework called the Attribution, Relation, and Order (ARO) benchmark to systematically measure how well vision-language models understand compositional relationships, object attributes, and word order. The authors evaluate popular foundation models—including CLIP, BLIP, FLAVA, and X-VLM—and examine why current training and evaluation practices mask severe compositional weaknesses.
To conduct this evaluation, the researchers constructed ARO using over 50,000 test cases derived from established visual datasets. The benchmark tests three core abilities: linking descriptive properties to the correct objects, identifying directed spatial and action relationships between entities, and recognizing correct sentence order compared to systematically scrambled versions. In parallel, the authors performed controlled experiments on standard image-text retrieval benchmarks by measuring model performance when captions were word-shuffled or when images were sliced and rearranged into mismatched visual patches.
The investigation produced four primary findings. First, widely used foundation models frequently operate like simple "bags of words," often performing near or below chance level when tasked with identifying proper relational direction (for instance, BLIP selected the inverted caption "the grass is eating the horse" with 81% probability). Second, models demonstrated minimal sensitivity to grammatical structure and word order, frequently scoring near the 20% random guessing baseline on word-order tasks. Third, conventional cross-modal retrieval benchmarks were shown to be flawed measures of comprehension; models suffered only marginal performance losses on standard retrieval tasks even when word order or image patches were completely scrambled. Fourth, the authors found that incorporating "composition-aware hard negatives" during training—specifically pairing scenes with close visual neighbors and injecting captions with swapped linguistic elements—dramatically closed this performance gap. Fine-tuning CLIP with this technique improved relational understanding accuracy from 59% to 81% and caption-order accuracy from 46% to 86% without degrading general image classification or retrieval capabilities.
These findings indicate that existing models achieve high benchmark scores through shortcut learning rather than genuine linguistic or visual understanding because standard retrieval pretraining does not require models to resolve fine-grained compositional differences. This reliance on lexical shortcuts poses practical operational risks, especially in downstream systems such as text-to-image generators, which can produce unfaithful outputs or reinforce societal stereotypes when attribute bindings fail. Organizations relying on off-the-shelf vision-language models for precise visual search, automated surveillance, or generative tasks face unexpected errors if models merely detect keyword presence rather than actual context.
To address these deficiencies, practitioners and developers should immediately integrate hard-negative mining strategies into model training pipelines and adopt fine-grained compositional benchmarks like ARO alongside standard metrics. Relying solely on conventional retrieval accuracy should be discontinued when evaluating readiness for deployment in mission-critical environments. Further research should prioritize pretraining full-scale foundation models with composition-aware objectives from scratch, as current empirical tests were limited to targeted fine-tuning on a single compute node with smaller batch sizes. Overall, the benchmark results provide high confidence that mainstream vision-language models have severe compositional blind spots, which can be substantially corrected with modest, targeted adjustments to contrastive training data.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Read the original CLIP paper first to understand the contrastive image–text training objective and model whose compositional limits the source evaluates.
- Paper: Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality, Anuj Diwan et al. (2022). This Winoground analysis establishes the compositionality failures that motivate ARO and clarifies why shared words in reordered captions pose a diagnostic challenge.
- Paper: VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena, Letitia Parcalabescu et al. (2022). VALSE provides a prior diagnostic benchmark for testing whether vision-language models ground linguistic phenomena, setting up the source’s broader tests of relations and order.
- Paper: Revisiting the Role of Language Priors in Vision-Language Models, Zhiqiu Lin et al. (2024). This later study revisits the source’s ARO compositionality problem and extends it by testing generative-model scoring and training-free correction of language-prior bias.
