Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality

Anuj DiwanLayne BerryEunsol ChoiDavid HarwathKyle Mahowald

article2022EMNLP56 citations

Reveals that vision-language model failures on the Winoground compositional benchmark stem primarily from cross-modal representation fusion and atypical visual reasoning demands rather than deficits in linguistic understanding.

Listen

Modern vision-and-language artificial intelligence models have achieved remarkable success on standard benchmarks such as image retrieval and visual question answering. However, these systems fail dramatically on the Winoground benchmark, performing no better than random chance when matching paired images with captions that share identical words in different orders (for example, distinguishing "a mug in some grass" from "some grass in a mug"). While prior explanations suggested that these models simply lack compositional linguistic understanding and rely only on word co-occurrences, the exact root causes of this breakdown have remained unclear.

The article investigates why vision-and-language models fail on Winoground by determining whether failures stem from limitations in language compositionality, visual perception, or the fusion of visual and textual information.

To evaluate these factors, the authors analyzed three representative multimodal architectures (CLIP, UNITER, and LXMERT) across multiple experimental settings. They relaxed the benchmark's strict zero-shot, two-candidate evaluation by measuring standard retrieval metrics across larger candidate pools and training classifier probes on model representations. They also established a fine-grained taxonomy of 400 Winoground items to isolate confounding challenges such as low visual quality, out-of-distribution content, and complex reasoning. Finally, they augmented the captions using nine natural language processing techniques to generate non-minimal paraphrases, evaluating whether separating the captions in linguistic embedding space enabled models to match them to the correct images.

The investigation produced four key findings. First, benchmark difficulty is not driven solely by lexical overlap: standard retrieval tests showed that LXMERT and UNITER struggle to match images and captions even when discriminating among unrelated items, though CLIP achieves over 78% retrieval accuracy in broader candidate pools. Second, manual categorization revealed that 229 of the 400 benchmark items involve challenges unrelated to core compositionality, such as visual blurriness, ambiguous descriptions, or complex reasoning; for example, CLIP achieved a 0% group accuracy on visually difficult items. Third, across 171 clean, vanilla compositionality items, all models still failed, with group scores remaining near random chance at 4% to 8%. Fourth, probing experiments demonstrated that the text branches of these models successfully separate caption meanings—with CLIP achieving over 80% accuracy in distinguishing text variants—yet incorporating these distinguishable variants failed to improve image-caption matching accuracy.

These results indicate that state-of-the-art vision-and-language models do possess sufficient linguistic capability to differentiate subtle textual meanings. The primary failure occurs when multimodal models attempt to fuse visual representations with text representations. Consequently, low performance on compositional benchmarks reflects integration bottlenecks rather than superficial language understanding alone.

Researchers and practitioners evaluating multimodal systems should report performance breakdowns using the article's fine-grained tags rather than relying solely on aggregate benchmark scores. Furthermore, development efforts aiming to improve compositional performance should prioritize cross-modal fusion mechanisms rather than focusing exclusively on language encoder pretraining.

The study's conclusions are bounded by its focus on three English-language transformer models and text-side data augmentations. While these findings demonstrate with high confidence that cross-modal representation alignment is a primary obstacle, visual-side augmentations and non-English linguistic structures require further exploration before generalizing across all multimodal domains.

No sufficiently relevant recommendations were found.

Cover for Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality

Abstract

Recent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning. Yet, they fail miserably on the recently proposed Winoground dataset (Thrush et al., 2022), which challenges models to match paired images and English captions, with items constructed to overlap lexically but differ in meaning (e.g., “there is a mug in some grass” vs. “there is some grass in a mug”). By annotating the dataset using new fine-grained tags, we show that solving the Winoground task requires not just compositional language understanding, but a host of other abilities like commonsense reasoning or locating small, out-of-focus objects in low-resolution images. In this paper, we identify the dataset’s main challenges through a suite of experiments on related tasks (probing task, image retrieval task), data augmentation, and manual inspection of the dataset. Our analysis suggests that a main challenge in visuolinguistic models may lie in fusing visual and textual representations, rather than in compositional language understanding. We release our annotation and code at https://github.com/ajd12342/why-winoground-hard.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Relaxing Winoground Constraints
  • 3.1 Recall at k > 1
  • 3.2 Task Adaptation
  • 4 Characterizing the Challenges Presented by Winoground Items
  • 4.1 Potentially Easy Pairs
  • 4.2 Potentially Difficult Pairs: In-Domain
  • 4.3 Potentially Difficult Pairs: Out-of-Domain
  • 4.4 Results on New Tags
  • 5 Generating Non-Minimal Winoground Data with Textual Variants
  • 5.1 Separability of Caption Variants
  • 5.2 Do Caption Variants Help with the Winoground Task?
  • 5.3 Distinguishing Captions Conditioned on Caption Variants
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Details on Evaluation Methods
  • A.1 Models
  • B Method for tagging
  • C Task Adaptation methods
  • D Textual Variants Methods
  • D.0.1 Syntactic augmentations
  • D.0.2 Semantic, word-based augmentations
  • D.0.3 Paraphrasing
  • D.0.4 Identity
  • D.1 Discriminable Caption Pair Experiment
  • E Winoground: New Tags

Knowls

  1. Knowl 1 — Winoground failures point to image–text alignment, not only caption understanding

    empirical result

    Across three vision–language models—CLIP, UNITER, and LXMERT—the study finds that Winoground failure is not explained simply by captions having high lexical overlap. LXMERT and UNITER text representations make many caption distinctions recoverable by probes, and a nonlinear probe can identify caption distinctions in CLIP representations with over 80% test accuracy. Yet making captions less lexically similar through augmentation does not materially improve matching them to images. The results therefore suggest that a central difficulty is connecting linguistic distinctions to visual representations, rather than an inability to represent caption meaning alone. This is an interpretation supported by the experiments, not a proof that fusion is the sole cause of failure.

  2. Knowl 2 — Six tags characterize non-compositional and additional Winoground challenges

    definition

    The authors annotate the 400 Winoground items with six non-exclusive tags, plus a NoTag category for items receiving none of the six. NonCompositional (30 items) marks lexical minimal pairs that are not semantic compositional variants, for example because a word belongs to an idiom, compound, or polysemous use. AmbiguouslyCorrect (46) marks items where at least one caption is correct for both images or neither when judged separately, so comparison between candidates is needed. VisuallyDifficult (38) marks items requiring detection of a small, blurry, background, out-of-focus, indistinct, or otherwise hard-to-see visual element. UnusualImage (56) marks items with an unrealistic or unusual image; UnusualText (50) marks captions that are misspelled, non-standard, ungrammatical, awkward, or difficult to parse. ComplexReasoning (78) marks items requiring commonsense or world knowledge, such as numerical or causal reasoning, non-English text understanding, or scientific knowledge. The remaining 171 items are NoTag. Annotators developed the categories by inspecting the dataset, applied them in a later pass, and had disagreements reviewed and resolved by consensus.

  3. Knowl 3 — Performance varies substantially across Winoground tag groups

    data/table

    For each item, Text score measures whether both captions are matched to their correct images, Image score whether both images are matched to their correct captions, and Group score whether both item-level scores are correct. Chance is 25% for Text and Image and 16.67% for Group. The following percentages are reported in the order Text / Image / Group; tag groups are not mutually exclusive.

    CLIP: NonCompositional 76.6 / 36.67 / 33.33; AmbiguouslyCorrect 30.43 / 15.22 / 13.04; VisuallyDifficult 15.79 / 0.00 / 0.00; UnusualImage 25.00 / 8.93 / 5.36; UnusualText 30.00 / 16.00 / 10.00; ComplexReasoning 24.36 / 7.69 / 3.85; NoTag 30.41 / 11.11 / 8.19.

    UNITER: NonCompositional 43.33 / 33.33 / 26.67; AmbiguouslyCorrect 30.43 / 13.04 / 8.70; VisuallyDifficult 31.58 / 7.89 / 5.26; UnusualImage 19.64 / 10.71 / 5.36; UnusualText 18.00 / 8.00 / 2.00; ComplexReasoning 29.49 / 6.41 / 3.85; NoTag 35.67 / 10.53 / 7.02.

    LXMERT: NonCompositional 10.00 / 13.33 / 3.33; AmbiguouslyCorrect 10.87 / 2.17 / 0.00; VisuallyDifficult 21.05 / 7.89 / 2.63; UnusualImage 12.50 / 7.14 / 1.79; UnusualText 10.00 / 6.00 / 0.00; ComplexReasoning 16.67 / 3.85 / 1.28; NoTag 19.88 / 4.68 / 4.09.

    The scores show that the easier NonCompositional group generally produces the strongest results, while CLIP scores zero on Image and Group for VisuallyDifficult items. The 171 untagged items also remain difficult, so removing items with the six identified challenges does not make the task easy.

  4. Knowl 4 — Winoground images and captions remain difficult in ordinary retrieval

    experimental setup

    The authors evaluated CLIP (ViT-B/32), UNITER-base, and LXMERT-base on retrieval using all 800 Winoground images and 800 captions. For each model, they scored all 800 × 800 image–caption combinations without fine-tuning. Recall at k (R@k) is the percentage of image or caption queries whose correct match appears among the top k ranked candidates; the paper reports text-to-image (T2I) and image-to-text (I2T) retrieval. R@1 requires ranking the correct match first, whereas larger k can succeed without resolving the within-item semantic minimal pair.

  5. Knowl 5 — Retrieval results expose large differences between model families

    data/table

    On the 800-image/800-caption retrieval evaluation, each pair below gives T2I / I2T recall percentages. For CLIP, R@1 was 32.9 / 27.4, R@2 54.4 / 47.9, R@5 72.4 / 65.9, and R@10 81.3 / 78.4. For UNITER, R@1 was 20.1 / 16.4, R@2 31.4 / 28.7, R@5 45.0 / 43.8, and R@10 55.3 / 55.4. For LXMERT, R@1 was 5.9 / 3.4, R@2 10.1 / 6.9, R@5 18.6 / 12.0, and R@10 26.5 / 15.6. CLIP performs substantially better than the other models at the less restrictive retrieval cutoffs, while LXMERT is weak at every cutoff. Thus, the models’ similar performance on the strict Winoground metric obscures markedly different retrieval ability on the same images and captions.

  6. Knowl 6 — Task-adaptation probes do not reliably recover Winoground matches

    experimental setup

    To test whether zero-shot evaluation or the original comparison format explains model failure, the authors trained probes to choose the better of two concatenated image–caption embeddings. They split the 400 items into 300 training and 100 test items, preserving tag proportions, and trained four-layer MLPs with hidden size 1,024 for 200 epochs on pooled-output embeddings from UNITER and LXMERT. Separate probes selected between two captions paired with one image or between two images paired with one caption. A control task used labels flipped for a random half of the items in each split. Each probe was run with 11 random seeds; the reported test results are minimum–maximum accuracy ranges.

  7. Knowl 7 — Adaptation probe accuracies remain near chance

    data/table

    For the task-adaptation probes, chance accuracy is 50%. The test-set ranges across 11 seeds were as follows. LXMERT target task: Text 49.0–54.5%, Image 48.5–51.8%; control: Text 42.2–58.5%, Image 48.2–57.2%. UNITER target task: Text 53.5–59.5%, Image 52.2–55.0%; control: Text 44.0–54.8%, Image 44.2–54.5%. No probe was appreciably and consistently better than chance or its control. This indicates that the pooled representations and probe setup used in this experiment did not reliably extract test-set matching information; it does not rule out every possible probe or representation.

  8. Knowl 8 — Caption variants do not improve image–caption matching

    empirical result

    The authors generated meaning-preserving caption variants using noun hyponym and hypernym replacements, synonym substitution, slang replacement, German-to-English backtranslation, two diverse-paraphrase methods, and semantic-role-label-based syntactic rewrites; the original caption was also included as an identity variant. They compared the original image–caption score with an aggregate over the original and variants, using either the mean or maximum aggregate and tuning the interpolation weight on group score. Mean aggregation with weight 0.5 worked best for LXMERT; maximum aggregation with weight 0.75 worked best for UNITER and CLIP.

    Original and augmented Text / Image / Group scores were: LXMERT 17.25 / 5.25 / 2.75 versus 17.50 / 4.75 / 3.25; UNITER 31.75 / 10.50 / 7.25 versus 31.50 / 12.50 / 8.25; CLIP 27.50 / 12.00 / 9.50 versus 27.25 / 12.25 / 9.75. Human original scores were 89.50 / 88.50 / 85.50; augmented human scores were not reported. The small, mixed changes show that increasing lexical variation in the caption candidates did not substantially resolve model matching failures.

  9. Knowl 9 — Caption distinctions are linearly accessible in LXMERT and UNITER embeddings

    empirical result

    For each Winoground item, the authors held the image input fixed and trained a linear support vector classifier to separate embeddings of variants of the two captions. A control classifier separated randomly assigned variant groups. The classifiers used CLS-token representations at model layers and C=100C=100; a successful separation means every variant was correctly classified, and margin width was measured as 2/∥w∥2/\lVert w\rVert, where ww is the classifier’s weight vector.

    Across layers and items, target-task classifiers perfectly separated caption groups in 81.3% of LXMERT cases, with mean margin width 1.9; the control succeeded in 10.9% of cases, with mean margin 0.7. For UNITER, target separation succeeded in 84.7% of cases, with mean margin 1.018, versus 19.0% and 0.48 for control. For CLIP, target separation succeeded in 3.5% of cases, with mean margin 0.13, while control never achieved perfect separation. The authors report that CLIP embeddings were not linearly separable after the first layer. These results indicate that caption distinctions are readily linearly accessible in LXMERT and UNITER representations, but not in CLIP’s under this linear-probe test.

  10. Knowl 10 — A nonlinear probe recovers caption distinctions from CLIP text representations

    empirical result

    A single four-layer MLP was trained across Winoground items to distinguish caption variants using text CLS embeddings, then evaluated on held-out items; unlike the per-item linear classifiers, this probe could exploit nonlinear patterns. CLIP’s target-task test accuracy exceeded 80% in the final layers, while its linear probes had rarely separated variants perfectly. LXMERT and UNITER target-task probe accuracy remained below 60%, although target probes outperformed randomized-label controls. For all three models, test accuracy rose in early layers and later declined around the introduction of cross-modal attention in LXMERT and in UNITER’s final two layers. The findings show that semantic caption distinctions can be recoverable from text representations—including nonlinearly for CLIP—even though this does not translate into successful image–caption matching.

  11. Knowl 11 — Scope of the conclusions is limited

    limitation

    The experiments cover only English and three multimodal models, so the authors do not claim that the findings generalize to other languages or all model architectures. The embedding-separability analysis focuses on text representations and does not include a parallel investigation of visual representations. The conclusions also rely partly on experiments that failed to improve performance; a different method could produce a different outcome.

Coverage note — The item-by-item tag assignments and detailed examples of individual augmentation outputs are omitted because they do not add a general result beyond the taxonomy, counts, procedures, and aggregate findings captured here.

References

  1. 1.Christoph Alt, Aleksandra Gabryszak, and Leonhard Hennig. 2020. Probing linguistic features of sentence-level representations in neural relation extraction. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1534–1545, Online. Association for Computational Linguistics.
  2. 2.Emily M. Bender and Alexander Koller. 2020. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online. Association for Computational Linguistics.
  3. 3.Steven Bird. 2006. NLTK: the natural language toolkit. In Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions, pages 69–72.
  4. 4.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8718–8735, Online. Association for Computational Linguistics.
  5. 5.Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz, and Michael Elhadad. 2021. Automatic generation of contrast sets from scene graphs: Probing the compositional consistency of GQA. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 94–105, Online. Association for Computational Linguistics.
  6. 6.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019. Uniter: Universal image-text representation learning.
  7. 7.Louis Clouatre, Philippe Trempe, Amal Zouaq, and Sarath Chandar. 2021. MLMLM: Link prediction with mean likelihood masked language model. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4321–4331, Online. Association for Computational Linguistics.
  8. 8.Kaustubh D. Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, Tongshuang Wu, Jascha Sohl-Dickstein, Jinho D. Choi, Eduard Hovy, Ondrej Dusek, Sebastian Ruder, Sajant Anand, Nagender Aneja, Rabin Banjade, Lisa Barthe, Hanna Behnke, Ian Berlot-Attwell, Connor Boyle, Caroline Brun, Marco Antonio Sobrevilla Cabezudo, Samuel Cahyawijaya, Emile Chapuis, Wanxiang Che, Mukund Choudhary, Christian Clauss, Pierre Colombo, Filip Cornell, Gautier Dagan, Mayukh Das, Tanay Dixit, Thomas Dopierre, Paul-Alexis Dray, Suchitra Dubey, Tatiana Ekeinhor, Marco Di Giovanni, Rishabh Gupta, Rishabh Gupta, Louanes Hamla, Sang Han, Fabrice Harel-Canada, Antoine Honore, Ishan Jindal, Przemyslaw K. Joniak, Denis Kleyko, Venelin Kovatchev, Kalpesh Krishna, Ashutosh Kumar, Stefan Langer, Seungjae Ryan Lee, Corey James Levinson, Hualou Liang, Kaizhao Liang, Zhexiong Liu, Andrey Lukyanenko, Vukosi Marivate, Gerard de Melo, Simon Meoni, Maxime Meyer, Afnan Mir, Nafise Sadat Moosavi, Niklas Muennighoff, Timothy Sum Hon Mun, Kenton Murray, Marcin Namysl, Maria Obedkova, Priti Oli, Nivranshu Pasricha, Jan Pfister, Richard Plant, Vinay Prabhu, Vasile Pais, Libo Qin, Shahab Raji, Pawan Kumar Rajpoot, Vikas Raunak, Roy Rinberg, Nicolas Roberts, Juan Diego Rodriguez, Claude Roux, Vasconcellos P. H. S., Ananya B. Sai, Robin M. Schmidt, Thomas Scialom, Tshephisho Sefara, Saqib N. Shamsi, Xudong Shen, Haoyue Shi, Yiwen Shi, Anna Shvets, Nick Siegel, Damien Sileo, Jamie Simon, Chandan Singh, Roman Sitelew, Priyank Soni, Taylor Sorensen, William Soto, Aman Srivastava, KV Aditya Srivatsa, Tony Sun, Mukund Varma T, A Tabassum, Fiona Anting Tan, Ryan Teehan, Mo Tiwari, Marie Tolkiehn, Athena Wang, Zijian Wang, Gloria Wang, Zijie J. Wang, Fuxuan Wei, Bryan Wilie, Genta Indra Winata, Xinyi Wu, Witold Wydmanski, Tianbao Xie, Usama Yaseen, M. Yee, Jing Zhang, and Yue Zhang. 2021. Nl-augmenter: A framework for task-sensitive natural language augmentation.
  9. 9.Thomas Dopierre, Christophe Gravier, and Wilfried Logerais. 2021. PROTAUGMENT: Unsupervised diverse short-texts paraphrasing for intent detection meta-learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2454–2466, Online. Association for Computational Linguistics.
  10. 10.Ashim Gupta, Giorgi Kvernadze, and Vivek Srikumar. 2021. BERT & family eat word salad: Experiments with text understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 12946–12954.
  11. 11.Jack Hessel and Alexandra Schofield. 2021. How effective is BERT without word ordering? implications for language understanding and data privacy. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 204–211, Online. Association for Computational Linguistics.
  12. 12.John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength Natural Language Processing in Python.
  14. 14.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR.
  15. 15.Najoung Kim and Tal Linzen. 2020. COGS: A compositional generalization challenge based on semantic interpretation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9087–9105, Online. Association for Computational Linguistics.
  16. 16.Ashutosh Kumar, Satwik Bhattamishra, Manik Bhandari, and Partha Talukdar. 2019. Submodular optimization-based diverse paraphrasing and its effectiveness in data augmentation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3609–3619, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The Winograd Schema Challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, KR’12, page 552–561. AAAI Press.
  18. 18.Alexandra Sasha Luccioni and David Rolnick. 2022. Bugs in the data: How imagenet misrepresents biodiversity.
  19. 19.Yiran Luo, Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral. 2022. To find waldo you need contextual cues: Debiasing who’s waldo. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 355–361, Dublin, Ireland. Association for Computational Linguistics.
  20. 20.Gary Marcus, Ernest Davis, and Scott Aaronson. 2022. A very preliminary analysis of DALL-E 2. arXiv preprint arXiv:2204.13807.
  21. 21.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  22. 22.George A Miller. 1998. WordNet: An electronic lexical database. MIT press.
  23. 23.Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019. Facebook FAIR’s WMT19 news translation task submission. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 314–319, Florence, Italy. Association for Computational Linguistics.
  24. 24.Joe O’Connor and Jacob Andreas. 2021. What context features can transformer language models use? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 851–864, Online. Association for Computational Linguistics.
  25. 25.Isabel Papadimitriou, Richard Futrell, and Kyle Mahowald. 2022. When classifying grammatical role, BERT doesn’t care about word order... except when it matters. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 636–643, Dublin, Ireland. Association for Computational Linguistics.
  26. 26.Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. 2021. Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1145–1160.
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision.
  28. 28.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  29. 29.Peng Shi and Jimmy Lin. 2019. Simple bert models for relation extraction and semantic role labeling.
  30. 30.Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021a. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2888–2913, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2021b. UnNatural Language Inference. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7329–7346, Online. Association for Computational Linguistics.
  32. 32.Paul Soulos, R. Thomas McCoy, Tal Linzen, and Paul Smolensky. 2020. Discovering the compositional structure of vector representations with role learning networks. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 238–254, Online. Association for Computational Linguistics.
  33. 33.Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017. A corpus of natural language for visual reasoning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 217–223, Vancouver, Canada. Association for Computational Linguistics.
  34. 34.Hao Tan and Mohit Bansal. 2019. LXMERT: learning cross-modality encoder representations from transformers. CoRR, abs/1908.07490.
  35. 35.Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022. Winoground: Probing vision and language models for visio-linguistic compositionality.
  36. 36.Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse beam search for improved description of complex scenes.

Citation

MLA
Diwan, A., et al. “Why Is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 2236–50, https://doi.org/10.18653/v1/2022.emnlp-main.143.
APA
Diwan, A., Berry, L., Choi, E., Harwath, D., & Mahowald, K. (2022). Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2236–2250. https://doi.org/10.18653/v1/2022.emnlp-main.143
Chicago
Diwan, A., L. Berry, E. Choi, D. Harwath, and K. Mahowald. 2022. “Why Is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2236–50. https://doi.org/10.18653/v1/2022.emnlp-main.143.
Harvard
Diwan, A. et al. (2022) “Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2236–2250. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.143.
Vancouver
1. Diwan A, Berry L, Choi E, Harwath D, Mahowald K (2022) Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2236–2250

BibTeX

@inproceedings{diwan-etal-2022-winoground,
    title = "Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality",
    author = "Diwan, Anuj  and
      Berry, Layne  and
      Choi, Eunsol  and
      Harwath, David  and
      Mahowald, Kyle",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.143/",
    doi = "10.18653/v1/2022.emnlp-main.143",
    pages = "2236--2250"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/