Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality
Anuj DiwanLayne BerryEunsol ChoiDavid HarwathKyle Mahowald
Reveals that vision-language model failures on the Winoground compositional benchmark stem primarily from cross-modal representation fusion and atypical visual reasoning demands rather than deficits in linguistic understanding.
Modern vision-and-language artificial intelligence models have achieved remarkable success on standard benchmarks such as image retrieval and visual question answering. However, these systems fail dramatically on the Winoground benchmark, performing no better than random chance when matching paired images with captions that share identical words in different orders (for example, distinguishing "a mug in some grass" from "some grass in a mug"). While prior explanations suggested that these models simply lack compositional linguistic understanding and rely only on word co-occurrences, the exact root causes of this breakdown have remained unclear.
The article investigates why vision-and-language models fail on Winoground by determining whether failures stem from limitations in language compositionality, visual perception, or the fusion of visual and textual information.
To evaluate these factors, the authors analyzed three representative multimodal architectures (CLIP, UNITER, and LXMERT) across multiple experimental settings. They relaxed the benchmark's strict zero-shot, two-candidate evaluation by measuring standard retrieval metrics across larger candidate pools and training classifier probes on model representations. They also established a fine-grained taxonomy of 400 Winoground items to isolate confounding challenges such as low visual quality, out-of-distribution content, and complex reasoning. Finally, they augmented the captions using nine natural language processing techniques to generate non-minimal paraphrases, evaluating whether separating the captions in linguistic embedding space enabled models to match them to the correct images.
The investigation produced four key findings. First, benchmark difficulty is not driven solely by lexical overlap: standard retrieval tests showed that LXMERT and UNITER struggle to match images and captions even when discriminating among unrelated items, though CLIP achieves over 78% retrieval accuracy in broader candidate pools. Second, manual categorization revealed that 229 of the 400 benchmark items involve challenges unrelated to core compositionality, such as visual blurriness, ambiguous descriptions, or complex reasoning; for example, CLIP achieved a 0% group accuracy on visually difficult items. Third, across 171 clean, vanilla compositionality items, all models still failed, with group scores remaining near random chance at 4% to 8%. Fourth, probing experiments demonstrated that the text branches of these models successfully separate caption meanings—with CLIP achieving over 80% accuracy in distinguishing text variants—yet incorporating these distinguishable variants failed to improve image-caption matching accuracy.
These results indicate that state-of-the-art vision-and-language models do possess sufficient linguistic capability to differentiate subtle textual meanings. The primary failure occurs when multimodal models attempt to fuse visual representations with text representations. Consequently, low performance on compositional benchmarks reflects integration bottlenecks rather than superficial language understanding alone.
Researchers and practitioners evaluating multimodal systems should report performance breakdowns using the article's fine-grained tags rather than relying solely on aggregate benchmark scores. Furthermore, development efforts aiming to improve compositional performance should prioritize cross-modal fusion mechanisms rather than focusing exclusively on language encoder pretraining.
The study's conclusions are bounded by its focus on three English-language transformer models and text-side data augmentations. While these findings demonstrate with high confidence that cross-modal representation alignment is a primary obstacle, visual-side augmentations and non-English linguistic structures require further exploration before generalizing across all multimodal domains.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). Read LXMERT first to understand one of the three evaluated architectures and its separate cross-modality encoder, whose Winoground failures the paper probes.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). UNITER is another model directly tested in the study, so its joint image-text representation and pretraining objectives clarify what the reported fusion failures mean.
No sufficiently relevant recommendations were found.
