Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Shengbang TongZhuang LiuYuexiang ZhaiYi MaYann LeCunSaining Xie
Reveals systematic visual failures in state-of-the-art multimodal models like GPT-4V caused by CLIP blind spots, introducing the Multimodal Visual Patterns benchmark to diagnose these errors and a Mixture of Features method combining self-supervised vision representations to correct them.
Recent advances in artificial intelligence have produced multimodal systems that combine large language models with visual inputs to describe scenes, answer questions, and follow complex user instructions. Despite impressive high-level reasoning, these models frequently fail at elementary visual perception tasks, such as recognizing object orientation, counting items, or identifying simple spatial relationships. These unexpected errors pose serious operational and reliability risks for deploying vision-language systems in mission-critical applications where accurate visual grounding is essential.
The article evaluates the root causes of these perceptual failures across leading multimodal systems and investigates whether standard visual encoders act as a performance bottleneck. To do this, the authors propose a targeted evaluation framework and test architectural modifications that integrate self-supervised visual features to improve visual accuracy.
The research identifies pairs of visually distinct images that standard contrastive vision-language encoders incorrectly map to nearly identical internal representations, termed blind pairs. Using these image pairs from standard image repositories, the authors built a benchmark containing 150 paired images and 300 unambiguous, straightforward questions spanning nine core visual patterns. They evaluated both proprietary systems (such as GPT-4V and Gemini) and leading open-source models (such as LLaVA and InstructBLIP), alongside a human baseline study. Additionally, they tested a feature-mixing strategy that merges representations from standard vision-language models with vision-only self-supervised models.
The findings show a substantial visual capability gap across all evaluated systems. Human participants answered 95.7% of the benchmark questions correctly, whereas the top-performing commercial models, Gemini and GPT-4V, achieved only 40.7% and 38.7% accuracy, respectively. Most open-source models scored below the 25% random-guessing baseline. The study revealed nine persistent visual patterns that challenge vision encoders, finding that seven of these categories cannot be solved by simply scaling up model parameters or training data. Furthermore, errors in the underlying vision encoder correlated strongly with overall model failure, showing Pearson correlation coefficients above 0.70 for open-source systems. Combining vision-only self-supervised features with contrastive features improved visual grounding accuracy on the benchmark by up to 10.7 percentage points without degrading general instruction-following capabilities.
These results demonstrate that language models are not solely responsible for multimodal errors; rather, commonly used vision encoders create an information bottleneck by overlooking granular visual details in favor of high-level concepts. Relying purely on standard classification metrics or scaling up existing models will not resolve these perceptual blind spots. Organizations deploying these systems should not assume strong language reasoning equates to dependable visual perception.
Decision-makers and developers should adopt multi-feature visual encoders that blend contrastive language-vision representations with self-supervised vision-only features to mitigate hallucinations and perception errors. Future development must also incorporate specialized visual benchmarks into evaluation pipelines rather than relying solely on traditional image classification metrics. While the findings provide strong evidence that visual bottlenecks exist, further testing is recommended on domain-specific workloads and varied operational resolutions before deploying these models in high-risk operational environments.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). Read MME first for its standardized evaluation of multimodal perception, which provides useful context for the source’s targeted tests of visual capability failures.
- Paper: VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena, Letitia Parcalabescu et al. (2022). VALSE’s diagnostic approach to testing visual grounding in counting and spatial relations helps situate the source’s focused benchmark of elementary perception.
- Paper: When and why vision-language models behave like bags-of-words, and what to do about it?, Mert Yuksekgonul et al. (2022). ARO establishes how vision-language models can miss object attributes and relations despite strong aggregate scores, framing the source’s diagnosis of failures hidden by broad evaluation.
No sufficiently relevant recommendations were found.
