Dense Connector for MLLMs
Huanjin YaoWenhao WuTaojiannan YangYuxin SongMengxi ZhangHaocheng FengYifan SunZhiheng LiWanli OuyangJingdong Wang
Proposes a plug-and-play vision-language connector that integrates multi-layer visual features from frozen vision encoders to boost multimodal large language model performance across 19 image and video benchmarks with minimal computational overhead.
Recent advances in artificial intelligence have produced multimodal models capable of processing both text and images, enabling capabilities across diverse applications. However, most research has focused on expanding language models or training datasets, largely treating visual inputs as fixed outputs from the final layer of a visual encoder. This common practice discards intermediate visual signals that capture different structural details and focus areas.
The article evaluates whether incorporating intermediate visual features from across a visual encoder's layers into the language model improves overall system performance. It demonstrates a plug-and-play architecture called Dense Connector, designed to enhance multimodal understanding with minimal computational overhead.
The authors conducted empirical evaluations across eleven image benchmarks and eight video question-answering benchmarks. The approach was tested across various model scales ranging from 2.7 billion to 70 billion parameters, different visual backbones (including CLIP and SigLIP), multiple image resolutions, and varying dataset sizes. The Dense Connector was implemented using three primary strategies: Sparse Token Integration, Sparse Channel Integration, and Dense Channel Integration, which aggregates features across grouped encoder layers.
The key findings indicate substantial performance and efficiency improvements. First, Dense Channel Integration consistently achieved state-of-the-art results across nineteen evaluation benchmarks, improving ScienceQA scores by 2.7 percentage points and MMBench scores by 2.5 percentage points over the standard baseline. Second, an efficient variant reduced the number of visual tokens by 75%—downsampling from 576 to 144 tokens—which cut fine-tuning runtime on eight high-end graphics processing units from 9 hours to 6.5 hours and yielded a three-fold increase in inference speed while outperforming baseline accuracy. Third, scaling to a 70-billion-parameter language model achieved a 97.8% score on visual instruction benchmarks, closely trailing proprietary commercial models. Finally, models trained purely on static images successfully transferred to video understanding without dedicated video training, achieving leading accuracies of 77.4% on MSVD-QA and 62.1% on MSRVTT-QA.
These findings suggest that organizations can achieve superior multimodal performance without the prohibitive financial and computational costs of training larger visual backbones from scratch. Harnessing intermediate features provides an effective upgrade path for existing workflows, and token reduction directly translates to lower operational costs and reduced latency in real-time deployments.
Organizations developing or deploying multimodal artificial intelligence should evaluate integrating multi-layer visual feature connectors into their model serving architectures. Engineering teams seeking inference cost reductions should prioritize token-efficient downsampling modules to accelerate throughput. Decision-makers should note that while results are strong across standard academic benchmarks, the authors note that attempts to add complex learnable parameters within the connector degraded training stability, and video processing occasionally mislabeled dynamic inputs as static scenes.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). It establishes the standard baseline visual instruction-tuning framework and projector designs that Dense Connector explicitly aims to improve upon.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). It introduces the fundamental paradigm of using lightweight linear projection connectors between frozen vision encoders and LLMs, which Dense Connector enhances with multi-layer intermediate features.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). It provides a foundational overview of multimodal large language model architectures, training stages, and connector interfaces.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). It details the architecture and alignment strategies of the open-source InternVL family that serve as core foundations and backbones for advanced vision-language connectors.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). It introduces visual alignment and projection techniques for video and image understanding that contextualize the video QA transfer results evaluated in Dense Connector.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It builds directly upon multimodal scaling insights by advancing model, data, and test-time scaling across the InternVL 2.5 series.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It extends open-source MLLM design by exploring native unified pre-training recipes and test-time scaling strategies in the InternVL3 architecture.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). It advances multimodal integration by incorporating cross-layer visual token injection mechanisms and spatial-temporal modeling in the Qwen3-VL framework.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). It continues the focus on visual token reduction by presenting a training-free filter-correlate-compress method to accelerate MLLM inference.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). It investigates visual token redundancy and downsampling strategies to optimize inference speed and latency in vision-language models.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It scales visual task transfer across single-image, multi-image, and video domains using unified representation and token budget constraints.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). It extends efficient token management by introducing hierarchical token compression tailored for high-resolution dynamic-cropping vision-language architectures.
