VCoder: Versatile Vision Encoders for Multimodal Large Language Models
Jitesh JainJianwei YangHumphrey Shi
Presents VCoder, a framework that integrates auxiliary perception modalities like segmentation and depth maps into multimodal large language models to resolve chronic errors in object identification, counting, and spatial hallucination.
Modern multimodal artificial intelligence systems perform exceptionally well at complex tasks such as visual question-answering and image captioning. However, they frequently fail at fundamental visual perception tasks, including identifying background objects, estimating depth order, and accurately counting items in cluttered scenes. This occurs largely because standard vision models focus predominantly on salient foreground objects described in caption datasets, leaving systems prone to significant hallucinations and basic perceptual errors.
The article demonstrates that augmenting Multimodal Large Language Models with dedicated perception inputs substantially improves their object identification, counting, and spatial ordering capabilities. Specifically, the article evaluates a new framework that feeds auxiliary perception signals directly into the language model architecture without disrupting its underlying reasoning capabilities.
To achieve this, the researchers developed an adapter framework called Versatile Vision Encoders, or VCoder, integrated into the open-source LLaVA-1.5 model. They constructed the COCO Segmentation Text dataset, which uses 280,000 training images paired with segmentation and depth maps from specialized off-the-shelf vision models to generate question-and-answer pairs covering semantic, instance, and panoptic object identification. The framework encodes these extra perception maps into distinct tokens, training only lightweight projection layers while freezing the base language and image components. The article also introduced standardized evaluation metrics: Count Score, Hallucination Score, and Depth Score.
The experimental findings show that VCoder significantly outperforms leading open-source models and proprietary systems like GPT-4V across all perception metrics. On panoptic object identification, VCoder achieved a Count Score of approximately 86% to 87% and reduced the Hallucination Score to around 11% to 12%, compared to GPT-4V, which achieved a Count Score of only 38.4% and an 83.0% Hallucination Score. Existing open-source models performed even lower, frequently recording Count Scores below 40%. Furthermore, for object depth ordering, VCoder reduced the Depth Score error from over 166 down to roughly 63 to 66.
These results indicate that relying purely on standard visual encoders introduces substantial operational risks for applications requiring precise inventory counting, spatial awareness, or environment tracking. Feeding dedicated perceptual representations directly into language models provides a computationally efficient path to eliminate object hallucinations without requiring expensive full-model retraining.
Organizations developing or deploying multimodal vision-language systems should incorporate dedicated perception adapters like VCoder rather than relying solely on raw image inputs for perception-critical tasks. Future technical efforts should focus on expanding the training dataset to broader, open-vocabulary object categories beyond standard closed sets and refining scoring metrics to handle vocabulary synonyms more flexibly without manual mapping.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Provides the foundational contrastive vision-language pre-training paradigm (CLIP) that underpins standard MLLM vision backbones and motivates the need for specialized perception encoders like VCoder.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Introduces the evaluation framework for measuring object hallucination and perception deficits in MLLMs, establishing the specific failure mode that VCoder targets.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). Establishes standard comprehensive benchmarks demonstrating that MLLMs systematically struggle with fine-grained visual perception tasks like object counting and localization.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Surveys the core architectural components and limitations of multimodal large language models, providing the essential context for integrating auxiliary perception encoders.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Presents foundational techniques for incorporating fine-grained visual grounding and localization capabilities into multimodal language models.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). Details methods for grounding visual regions into discrete spatial tokens within causal language models, a key reference point for spatial perception in MLLMs.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). Demonstrates the canonical projector-based architecture connecting frozen visual encoders to LLMs, which VCoder builds upon by introducing versatile perception encoders.
- Paper: OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding, Tao Zhang et al. (2024). Extends the integration of perception modalities into MLLMs by unifying image-, object-, and pixel-level visual prompting and reasoning in a single framework.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). Systematically explores the design space of multi-encoder vision representations, analyzing the fusion of low-level spatial features and high-level semantics in MLLMs.
- Paper: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention, Wenbin An et al. (2025). Builds upon object perception and hallucination challenges by introducing attention assembly mechanisms to focus MLLMs on fine-grained visual regions during decoding.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). Continues the effort to improve object grounding and factual perception in MLLMs via visual information grounding and mutual-information decoding.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). Investigates whether enhanced spatial and perception mechanisms in MLLMs generalize to understanding and reasoning over complex 3D environments.
