Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Shengbang TongZhuang LiuYuexiang ZhaiYi MaYann LeCunSaining Xie

article2024CVPR681 citations

Reveals systematic visual failures in state-of-the-art multimodal models like GPT-4V caused by CLIP blind spots, introducing the Multimodal Visual Patterns benchmark to diagnose these errors and a Mixture of Features method combining self-supervised vision representations to correct them.

Listen

Recent advances in artificial intelligence have produced multimodal systems that combine large language models with visual inputs to describe scenes, answer questions, and follow complex user instructions. Despite impressive high-level reasoning, these models frequently fail at elementary visual perception tasks, such as recognizing object orientation, counting items, or identifying simple spatial relationships. These unexpected errors pose serious operational and reliability risks for deploying vision-language systems in mission-critical applications where accurate visual grounding is essential.

The article evaluates the root causes of these perceptual failures across leading multimodal systems and investigates whether standard visual encoders act as a performance bottleneck. To do this, the authors propose a targeted evaluation framework and test architectural modifications that integrate self-supervised visual features to improve visual accuracy.

The research identifies pairs of visually distinct images that standard contrastive vision-language encoders incorrectly map to nearly identical internal representations, termed blind pairs. Using these image pairs from standard image repositories, the authors built a benchmark containing 150 paired images and 300 unambiguous, straightforward questions spanning nine core visual patterns. They evaluated both proprietary systems (such as GPT-4V and Gemini) and leading open-source models (such as LLaVA and InstructBLIP), alongside a human baseline study. Additionally, they tested a feature-mixing strategy that merges representations from standard vision-language models with vision-only self-supervised models.

The findings show a substantial visual capability gap across all evaluated systems. Human participants answered 95.7% of the benchmark questions correctly, whereas the top-performing commercial models, Gemini and GPT-4V, achieved only 40.7% and 38.7% accuracy, respectively. Most open-source models scored below the 25% random-guessing baseline. The study revealed nine persistent visual patterns that challenge vision encoders, finding that seven of these categories cannot be solved by simply scaling up model parameters or training data. Furthermore, errors in the underlying vision encoder correlated strongly with overall model failure, showing Pearson correlation coefficients above 0.70 for open-source systems. Combining vision-only self-supervised features with contrastive features improved visual grounding accuracy on the benchmark by up to 10.7 percentage points without degrading general instruction-following capabilities.

These results demonstrate that language models are not solely responsible for multimodal errors; rather, commonly used vision encoders create an information bottleneck by overlooking granular visual details in favor of high-level concepts. Relying purely on standard classification metrics or scaling up existing models will not resolve these perceptual blind spots. Organizations deploying these systems should not assume strong language reasoning equates to dependable visual perception.

Decision-makers and developers should adopt multi-feature visual encoders that blend contrastive language-vision representations with self-supervised vision-only features to mitigate hallucinations and perception errors. Future development must also incorporate specialized visual benchmarks into evaluation pipelines rather than relying solely on traditional image classification metrics. While the findings provide strong evidence that visual bottlenecks exist, further testing is recommended on domain-specific workloads and varied operational resolutions before deploying these models in high-risk operational environments.

No sufficiently relevant recommendations were found.

Cover for Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Abstract

Is vision good enough for language? Recent advancements in multimodal models primarily stem from the powerful reasoning abilities of large language models (LLMs). However, the visual component typically depends only on the instance-level contrastive language-image pre-training (CLIP). Our research reveals that the visual capabilities in recent multimodal LLMs (MLLMs) still exhibit systematic shortcomings. To understand the roots of these errors, we explore the gap between the visual embedding space of CLIP and vision-only self-supervised learning. We identify ''CLIP-blind pairs'' - images that CLIP perceives as similar despite their clear visual differences. With these pairs, we construct the Multimodal Visual Patterns (MMVP) benchmark. MMVP exposes areas where state-of-the-art systems, including GPT-4V, struggle with straightforward questions across nine basic visual patterns, often providing incorrect answers and hallucinated explanations. We further evaluate various CLIP-based vision-and-language models and found a notable correlation between visual patterns that challenge CLIP models and those problematic for multimodal LLMs. As an initial effort to address these issues, we propose a Mixture of Features (MoF) approach, demonstrating that integrating vision self-supervised learning features with MLLMs can significantly enhance their visual grounding capabilities. Together, our research suggests visual representation learning remains an open challenge, and accurate visual grounding is crucial for future successful multimodal systems.

Table of Contents

  • 1 Introduction
  • 2 The Multimodal Visual Patterns (MMVP) Benchmark
  • 2.1 Finding CLIP-blind Pairs
  • 2.2 Designing Benchmark from CLIP-blind Pairs
  • 2.3 Benchmark Results
  • 3 Systematic Failures in CLIP
  • 3.1 Visual Patterns in CLIP-blind Pairs
  • 3.2 The MMVP-VLM Benchmark
  • 3.3 How CLIP’s Errors Affect MLLMs
  • 4 Mixture-of-Features (MoF) for MLLM
  • 4.1 Experiment Setting
  • 4.2 Additive MoF
  • 4.3 Interleaved MoF
  • 5 Related Works
  • 6 Discussion
  • References
  • A Experiment Details
  • B MMVP Benchmark
  • B.1 Details of evaluating SOTA models
  • B.2 Questions in MMVP Benchmark
  • B.3 Ablation Studies
  • B.4 Human Study Details
  • C CLIP-MLLM Failure Correlation
  • D Visual Patterns for CLIP
  • E More Benchmark Results
  • E.1 Different vision-only backbones
  • E.2 Scaling up to larger resolution

Knowls

  1. Knowl 1 — Operational definition of CLIP-blind image pairs

    definition

    A CLIP-blind pair is a pair of images whose embeddings are highly similar under a CLIP vision encoder but dissimilar under a vision-only self-supervised encoder. The paper searches ImageNet and LAION-Aesthetics using CLIP-ViT-L-14 and DINOv2-ViT-L-14, retaining pairs with cosine similarity greater than 0.950.95 in CLIP space and less than 0.60.6 in DINOv2 space. The authors use DINOv2 distance as an indicator that the images differ visually, and hypothesize that CLIP has encoded at least one image ambiguously when the two encoders disagree in this way.

  2. Knowl 2 — Construction and scoring of the MMVP benchmark

    experimental setup

    The Multimodal Visual Patterns (MMVP) benchmark contains 150 CLIP-blind image pairs and 300 manually written visual-question-answering questions. For each pair, annotators identify a visual difference and write straightforward, unambiguous questions about the distinguishing details in the two images. A model is counted as correct on a pair only if it answers the questions for both images correctly; getting just one image right earns no pair-level credit. This paired scoring tests whether a model can distinguish details that CLIP embeddings treat as similar.

  3. Knowl 3 — MLLM performance on MMVP versus human performance

    empirical result

    On MMVP, reported accuracy was 95.7% for humans, compared with 40.7% for Gemini and 38.7% for GPT-4V. The remaining tested systems scored 24.7% for LLaVA-1.5, 19.0% for Bard, 17.3% for Bing Chat, 16.7% for InstructBLIP, 12.7% for mini-GPT4, and 6.0% for LLaVA; the reported random-guess level was 25.0%. Thus, all tested systems except GPT-4V and Gemini were below that reference level, and the strongest models remained far below human performance. The human result is the average from a four-person study; model questions were queried independently to avoid chat-history effects.

  4. Knowl 4 — Nine visual patterns used to characterize CLIP failures

    definition

    The paper groups the visual distinctions probed by MMVP into nine patterns: (1) orientation and direction, such as which way an animal faces; (2) presence of specific features, whether an element is present or absent; (3) state and condition, such as whether an object is open or closed; (4) quantity and count; (5) positional and relational context between objects; (6) color and appearance; (7) structural and physical characteristics; (8) text or symbols in an image; and (9) viewpoint and perspective, meaning the angle from which a scene or object is viewed. GPT-4 was prompted to help categorize the benchmark questions into these general visual patterns.

  5. Knowl 5 — MMVP-VLM tests whether CLIP scaling resolves visual-pattern failures

    empirical result

    MMVP-VLM evaluates image-text matching across the nine visual patterns, with 15 text-image pairs per pattern; a pair counts as correct only when both image-text matches are correct. Across the tested CLIP-family models, average accuracy ranged from 19.3% to 39.3%. The highest average was 39.3% for DFN ViT-H-14 at image size 224, followed by 37.8% for SigLIP ViT-SO-14 at 224; the corresponding scores for larger image sizes were 34.8% for DFN at 378 and 37.0% for SigLIP at 384. The authors report that scaling model size or training data helped only on state-and-condition and color-and-appearance patterns; the other seven remained difficult across the tested models. Increasing resolution alone yielded little improvement. For example, the best reported scores across models were 26.7% for orientation, 26.7% for presence of specific features, 73.3% for state and condition, 40.0% for quantity and count, 33.3% for positional and relational context, 66.7% for color and appearance, 46.7% for structural characteristics, 26.7% for text, and 53.3% for viewpoint and perspective. ImageNet-1k zero-shot accuracy, which ranged from 75.5% to 84.4% in these experiments, did not reliably predict MMVP-VLM performance, especially among the stronger ImageNet models.

  6. Knowl 6 — CLIP and MLLM errors correlate across visual patterns

    empirical result

    The paper compares accuracy by visual pattern for CLIP and several MLLMs, then calculates Pearson correlations across the nine patterns. The reported coefficients are 0.87 for LLaVA-1.5, 0.71 for InstructBLIP, 0.79 for Bard, 0.72 for Gemini, and 0.31 for GPT-4. The authors interpret the strong correlations for LLaVA-1.5 and InstructBLIP, which explicitly use CLIP-based visual encoders, as evidence that CLIP's pattern-specific weaknesses are reflected in downstream MLLM performance. The correlations are associations across pattern-level scores, not evidence that the encoder alone causes every MLLM error.

  7. Knowl 7 — Additive Mixture-of-Features combines CLIP and DINOv2 features

    model/method

    Additive Mixture-of-Features (A-MoF) combines image features from a pretrained CLIP encoder and a pretrained DINOv2 vision-only self-supervised encoder before the LLaVA adapter. If zC\mathbf{z}_{C} is the CLIP feature and zD\mathbf{z}_{D} is the DINOv2 feature, the mixed feature is zmix=αzC+(1−α)zD\mathbf{z}_{mix}=\alpha\mathbf{z}_{C}+(1-\alpha)\mathbf{z}_{D}, where α\alpha controls the CLIP contribution and 1−α1-\alpha the DINOv2 contribution. The experiments vary the DINOv2 proportion over 00, 0.250.25, 0.500.50, 0.6250.625, 0.750.75, 0.8750.875, and 1.01.0. The approach uses the LLaVA framework and the CLIP-ViT-L-14 and DINOv2-ViT-L-14 encoders.

  8. Knowl 8 — A-MoF trades instruction following for visual grounding

    empirical result

    In the A-MoF experiments, the baseline LLaVA model with no DINOv2 contribution scored 5.5 on MMVP and 81.8 on the LLaVA instruction-following benchmark. As the DINOv2 proportion increased, the results were: proportion 0.25, MMVP 7.9 and LLaVA 79.4; 0.50, 12.0 and 78.6; 0.625, 15.0 and 76.4; 0.75, 18.7 and 75.8; 0.875, 16.5 and 69.3; and 1.0, 13.4 and 68.5. The strongest MMVP score occurred at a DINOv2 proportion of 0.75, while instruction-following scores declined as DINOv2 contribution rose, with a particularly marked decline at 0.875. The results therefore show a trade-off for additive mixing rather than a uniform improvement on both capabilities.

  9. Knowl 9 — Interleaved Mixture-of-Features preserves spatial ordering

    model/method

    Interleaved Mixture-of-Features (I-MoF) sends the same image through CLIP and a vision-only self-supervised encoder, processes each encoder's output through its own adapter, and then interleaves the resulting visual tokens for the language model while retaining their spatial ordering. Unlike A-MoF, I-MoF does not linearly blend the two feature sets into one representation before the adapter; it supplies tokens from both representations to the LLM. The paper instantiates this approach in LLaVA with CLIP and DINOv2, and also reports experiments using MAE and MoCoV3 as the self-supervised encoder.

  10. Knowl 10 — I-MoF improves grounding across models and benchmarks

    empirical result

    With LLaVA at image size 224, I-MoF raised MMVP from 5.5 to 16.7, while the LLaVA benchmark score changed from 81.8 to 82.8 and POPE from 50.0 to 51.0. Raising baseline LLaVA resolution from 224 to 336 changed MMVP only from 5.5 to 6.0. With LLaVA-1.5, the 224-resolution I-MoF model scored 28.0 on MMVP and 86.3 on POPE, compared with 24.7 and 85.9 for the 336-resolution baseline; its LLaVA benchmark score was 82.7 versus 84.7. At 336 resolution, LLaVA-1.5 with I-MoF reached 31.3 on MMVP, 86.7 on POPE, and 81.8 on the LLaVA benchmark. The authors also found MMVP gains with other self-supervised encoders: for LLaVA-1.5, MoCoV3 I-MoF scored 26.7 on MMVP and 86.1 on POPE, MAE I-MoF scored 27.3 and 86.1, and DINOv2 I-MoF scored 28.0 and 86.3; the baseline scored 24.7 and 85.9. In the additional 336-resolution comparison, I-MoF scored 81.8 on the LLaVA benchmark, 73.3 on LLaVA-In-the-Wild, 65.4 on MMBench, 58.7 on TextVQA, 86.7 on POPE, 79.3 on VQA-v2, and 34.6 on MM-Vet, versus baseline scores of 84.7, 70.7, 67.7, 61.3, 85.9, 80.0, and 35.4, respectively. These experiments support the paper's claim that interleaving can improve visual grounding while keeping performance on general instruction and multimodal benchmarks broadly comparable, though not higher on every measure.

Coverage note — The full set of 300 MMVP questions and detailed training-hyperparameter tables are omitted because they provide item-level examples or implementation settings rather than additional standalone findings; the benchmark construction, principal evaluation results, and MoF results are included.

References

  1. 1.ShareGPT, 2023.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In NeruIPS, 2022.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  4. 4.Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In CVPR, 2023.
  5. 5.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. 2022.
  6. 6.Sebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  7. 7.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020.
  8. 8.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICML, 2021.
  10. 10.Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023.
  11. 11.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  12. 12.Hila Gonen and Yoav Goldberg. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In NAACL, 2019.
  13. 13.Google. Bard, 2023.
  14. 14.Google. Gemini, 2023.
  15. 15.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017.
  16. 16.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. In NeurIPS, 2020.
  17. 17.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  18. 18.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.
  19. 19.Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality. In NeurIPS, 2023.
  20. 20.Jennifer Hu and Roger Levy. Prompt-based methods may underestimate large language models’ linguistic generalizations. In EMNLP, 2023.
  21. 21.Drew A Hudson and Christopher D Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019.
  22. 22.Heewoo Jun and Alex Nichol. Shap-E: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023.
  23. 23.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014.
  24. 24.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalanditis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017.
  25. 25.Hugo Laurenc¸on, Lucile Saulnier, Leo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M Rush, Douwe Kiela, et al. Obelisc: An open web-scale filtered dataset of interleaved image-text documents. arXiv preprint arXiv:2306.16527, 2023.
  26. 26.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023.
  27. 27.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023.
  28. 28.Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for GPT-4V (ision), LLaVA-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566, 2023.
  29. 29.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023.
  30. 30.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023.
  31. 31.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. 2023.
  32. 32.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023.
  33. 33.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2017.
  34. 34.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In CVPR, 2016.
  35. 35.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR, 2019.
  36. 36.Chandler May, Alex Wang, Shikha Bordia, Samuel R Bowman, and Rachel Rudinger. On measuring social biases in sentence encoders. In NAACL, 2019.
  37. 37.Microsoft. newbing, 2023.
  38. 38.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. OCR-VQA: Visual question answering by reading text in images. In ICDAR, 2019.
  39. 39.Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie. Slip: Self-supervision meets language-image pre-training. In ECCV, 2022.
  40. 40.OpenAI. GPT-4V(ision) System Card, 2023.
  41. 41.OpenAI. Gpt-4 technical report, 2023.
  42. 42.Maxime Oquab, Timothee Darcet, Th´eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  43. 43.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020.
  45. 45.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. In NeurIPS, 2021.
  46. 46.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  47. 47.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015.
  48. 48.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: An open large-scale dataset for training next generation image-text models. In NeurIPS, 2022.
  49. 49.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. In ECCV, 2022.
  50. 50.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018.
  51. 51.Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: a dataset for image captioning with reading comprehension. In ECCV, 2020.
  52. 52.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In CVPR, 2019.
  53. 53.Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan, Vaibhav Aggarwal, Aaron Adcock, Armand Joulin, Piotr Dollar, Christoph Feichtenhofer, Ross Girshick, et al. The effectiveness of MAE pre-pretraining for billion-scale pretraining. In ICCV, 2023.
  54. 54.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. EVA-CLIP: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  55. 55.Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. Mitigating gender bias in natural language processing: Literature review. In ACL, 2019.
  56. 56.Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing vision and language models for visio-linguistic compositionality. In CVPR, 2022.
  57. 57.Shengbang Tong, Erik Jones, and Jacob Steinhardt. Mass-producing failures of multimodal systems with language models. In NeurIPS, 2023.
  58. 58.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  59. 59.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. LLaMA 2: Open foundation and fine-tuned chat models. 2023.
  60. 60.Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, and Lucas Beyer. Image captioners are scalable vision learners too. NeurIPS, 2023.
  61. 61.Kirill Vishniakov, Zhiqiang Shen, and Zhuang Liu. Convnet vs transformer, supervised vs clip: Beyond imagenet accuracy, 2024.
  62. 62.Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer. Demystifying CLIP data. arXiv preprint arXiv:2309.16671, 2023.
  63. 63.Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs: Preliminary Explorations with GPT-4V (ision). arXiv preprint arXiv:2309.17421, 2023.
  64. 64.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. MM-Vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  65. 65.Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and why vision-language models behave like bags-of-words, and what to do about it? In ICLR, 2022.
  66. 66.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023.
  67. 67.Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. Investigating the catastrophic forgetting in multimodal large language models. arXiv preprint arXiv:2309.10313, 2023.
  68. 68.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023.
  69. 69.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena, 2023.
  70. 70.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. iBOT: Image BERT pre-training with online tokenizer. In ICLR, 2021.
  71. 71.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Tong, S., et al. “Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs”. arXiv, 2024, http://arxiv.org/abs/2401.06209v2.
APA
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., & Xie, S. (2024). Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. arXiv. http://arxiv.org/abs/2401.06209v2
Chicago
Tong, S., Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie. 2024. “Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs”. arXiv. http://arxiv.org/abs/2401.06209v2.
Harvard
Tong, S. et al. (2024) “Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.06209v2.
Vancouver
1. Tong S, Liu Z, Zhai Y, Ma Y, LeCun Y, Xie S (2024) Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. arXiv

BibTeX

@article{tong2024eyes,
  title = {Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs},
  author = {Tong, Shengbang and Liu, Zhuang and Zhai, Yuexiang and Ma, Yi and LeCun, Yann and Xie, Saining},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.06209v2},
  eprint = {2401.06209}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE