Dense Connector for MLLMs

Huanjin YaoWenhao WuTaojiannan YangYuxin SongMengxi ZhangHaocheng FengYifan SunZhiheng LiWanli OuyangJingdong Wang

article2024NeurIPS64 citations

Proposes a plug-and-play vision-language connector that integrates multi-layer visual features from frozen vision encoders to boost multimodal large language model performance across 19 image and video benchmarks with minimal computational overhead.

Listen

Recent advances in artificial intelligence have produced multimodal models capable of processing both text and images, enabling capabilities across diverse applications. However, most research has focused on expanding language models or training datasets, largely treating visual inputs as fixed outputs from the final layer of a visual encoder. This common practice discards intermediate visual signals that capture different structural details and focus areas.

The article evaluates whether incorporating intermediate visual features from across a visual encoder's layers into the language model improves overall system performance. It demonstrates a plug-and-play architecture called Dense Connector, designed to enhance multimodal understanding with minimal computational overhead.

The authors conducted empirical evaluations across eleven image benchmarks and eight video question-answering benchmarks. The approach was tested across various model scales ranging from 2.7 billion to 70 billion parameters, different visual backbones (including CLIP and SigLIP), multiple image resolutions, and varying dataset sizes. The Dense Connector was implemented using three primary strategies: Sparse Token Integration, Sparse Channel Integration, and Dense Channel Integration, which aggregates features across grouped encoder layers.

The key findings indicate substantial performance and efficiency improvements. First, Dense Channel Integration consistently achieved state-of-the-art results across nineteen evaluation benchmarks, improving ScienceQA scores by 2.7 percentage points and MMBench scores by 2.5 percentage points over the standard baseline. Second, an efficient variant reduced the number of visual tokens by 75%—downsampling from 576 to 144 tokens—which cut fine-tuning runtime on eight high-end graphics processing units from 9 hours to 6.5 hours and yielded a three-fold increase in inference speed while outperforming baseline accuracy. Third, scaling to a 70-billion-parameter language model achieved a 97.8% score on visual instruction benchmarks, closely trailing proprietary commercial models. Finally, models trained purely on static images successfully transferred to video understanding without dedicated video training, achieving leading accuracies of 77.4% on MSVD-QA and 62.1% on MSRVTT-QA.

These findings suggest that organizations can achieve superior multimodal performance without the prohibitive financial and computational costs of training larger visual backbones from scratch. Harnessing intermediate features provides an effective upgrade path for existing workflows, and token reduction directly translates to lower operational costs and reduced latency in real-time deployments.

Organizations developing or deploying multimodal artificial intelligence should evaluate integrating multi-layer visual feature connectors into their model serving architectures. Engineering teams seeking inference cost reductions should prioritize token-efficient downsampling modules to accelerate throughput. Decision-makers should note that while results are strong across standard academic benchmarks, the authors note that attempts to add complex learnable parameters within the connector degraded training stability, and video processing occasionally mislabeled dynamic inputs as static scenes.

arXiv: 2405.13800
Cover for Dense Connector for MLLMs

Abstract

Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garnered broad attention from both academia and industry. In the current MLLM rat race, the focus seems to be predominantly on the linguistic side. We witness the rise of larger and higher-quality instruction datasets, as well as the involvement of larger-sized LLMs. Yet, scant attention has been directed towards the visual signals utilized by MLLMs, often assumed to be the final high-level features extracted by a frozen visual encoder. In this paper, we introduce the Dense Connector - a simple, effective, and plug-and-play vision-language connector that significantly enhances existing MLLMs by leveraging multi-layer visual features, with minimal additional computational overhead. Building on this, we also propose the Efficient Dense Connector, which achieves performance comparable to LLaVA-v1.5 with only 25% of the visual tokens. Furthermore, our model, trained solely on images, showcases remarkable zero-shot capabilities in video understanding as well. Experimental results across various vision encoders, image resolutions, training dataset scales, varying sizes of LLMs (2.7B→70B), and diverse architectures of MLLMs (e.g., LLaVA-v1.5, LLaVA-NeXT, and Mini-Gemini) validate the versatility and scalability of our approach, achieving state-of-the-art performance across 19 image and video benchmarks. We hope that this work will provide valuable experience and serve as a basic module for future MLLM development. Code is available at https://github.com/HJYao00/DenseConnector.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Pre-trained Vision Models
  • 2.2 Large Language Models
  • 2.3 Multimodal Large Language Models
  • 3 Method
  • 3.1 Overview
  • 3.2 Dense Connector
  • 3.3 Efficient Dense Connector for Visual Token Optimization
  • 3.4 Training-Free Extension from Image to Video Conversational Models
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Ablation Study
  • 4.3 Main Results
  • 5 Conclusion and Limitation
  • References
  • A Appendix
  • A.1 Further Exploration and Analysis of the Dense Connector
  • A.2 Visualized Analysis
  • A.3 Model Zoo
  • A.4 Further Exploration in Training Visual Encoder
  • A.5 More Qualitative Results
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Dense Channel Integration (DCI) in Dense Connector

    model/method

    The Dense Connector (DC) integrates multi-layer visual representations from a frozen Vision Transformer (ViT) into a Multimodal Large Language Model (MLLM) without increasing the computational cost of the visual encoder. Let an input image Xi∈RH×W×CX_i \in \mathbb{R}^{H \times W \times C} be processed by a ViT of LL layers to yield multi-layer visual features V=[V1,…,VL]∈RL×N×DvV = [V_1, \dots, V_L] \in \mathbb{R}^{L \times N \times D_v}, where NN is the number of patch tokens and DvD_v is the visual hidden dimension.

    In Dense Channel Integration (DCI), all LL layers of features are partitioned into GG contiguous groups, where each group contains M=L/GM = L/G adjacent layer features. The features within each group g∈{1,…,G}g \in \{1, \dots, G\} are aggregated via an element-wise arithmetic average:

    GVg=1M∑i=(g−1)M+1gMVi,1≤g≤GGV_g = \frac{1}{M} \sum_{i=(g-1)M+1}^{gM} V_i, \quad 1 \le g \le G

    The GG grouped representations are concatenated with the top (final) layer features VL∈RN×DvV_L \in \mathbb{R}^{N \times D_v} along the channel dimension, resulting in a feature tensor of size N×(G+1)DvN \times (G+1)D_v. This concatenated representation is then passed through a learnable Multi-Layer Perceptron (MLP) consisting of two linear layers with an intermediate Gaussian Error Linear Unit (GELU) activation function to project the visual embeddings into the language model's hidden dimension DtD_t:

    ev=MLP(Concatenate([GV1,…,GVG,VL],dim=channel))∈RN×Dte_v = \text{MLP}(\text{Concatenate}([GV_1, \dots, GV_G, V_L], \text{dim}=\text{channel})) \in \mathbb{R}^{N \times D_t}

    By default, for a 24-layer ViT (e.g., CLIP-ViT-L), the layers are divided into G=2G=2 groups (M=12M=12 layers per group: layers 1–12 and 13–24). For a 27-layer vision encoder (e.g., SigLIP-ViT-SO), the first 26 layers are partitioned into 2 groups (layers 1–13 and 14–26). DCI preserves the spatial token length NN, avoiding increased sequence processing overhead in the LLM.

  2. Knowl 2 — Sparse Token Integration (STI) and Sparse Channel Integration (SCI)

    model/method

    Sparse Token Integration (STI) and Sparse Channel Integration (SCI) are two alternative instantiations of the Dense Connector that incorporate intermediate representations from KK selected layers {l1,…,lK}\{l_1, \dots, l_K\} (where 1≤ln<L1 \le l_n < L) alongside the final layer LL of an LL-layer Vision Transformer producing token features V∈RL×N×DvV \in \mathbb{R}^{L \times N \times D_v}.

    1. Sparse Token Integration (STI): Intermediate layer features VlnV_{l_n} are spatially downsampled using 2D average pooling avg(⋅)\text{avg}(\cdot) with stride α\alpha to reduce the token count from NN to N′=N/αN' = N/\alpha. The final layer visual features VL∈RN×DvV_L \in \mathbb{R}^{N \times D_v} are preserved at full resolution. These features are concatenated along the token sequence dimension and mapped into the LLM hidden space DtD_t via a shared 2-layer MLP projector:

    ev=MLP(Concatenate([avg(Vl1),…,avg(VlK),VL],dim=token))∈R(N+K⋅N′)×Dte_v = \text{MLP}(\text{Concatenate}([\text{avg}(V_{l_1}), \dots, \text{avg}(V_{l_K}), V_L], \text{dim}=\text{token})) \in \mathbb{R}^{(N + K \cdot N') \times D_t}

    1. Sparse Channel Integration (SCI): All selected intermediate layer features and the final layer features are preserved at full token resolution NN and concatenated along the channel dimension. The 2-layer MLP projector maps the combined (K+1)Dv(K+1)D_v-dimensional channels down to the LLM hidden dimension DtD_t, keeping the number of visual tokens fixed at NN:

    ev=MLP(Concatenate([Vl1,…,VlK,VL],dim=channel))∈RN×Dte_v = \text{MLP}(\text{Concatenate}([V_{l_1}, \dots, V_{l_K}, V_L], \text{dim}=\text{channel})) \in \mathbb{R}^{N \times D_t}

  3. Knowl 3 — Efficient Dense Connector for Visual Token Reduction

    model/method

    The Efficient Dense Connector reduces visual sequence length in autoregressive MLLMs by applying a parameter-free spatial downsampling operation immediately following the Dense Connector projection.

    After generating the multi-layer visual token representation ev∈RN×Dte_v \in \mathbb{R}^{N \times D_t} through the Dense Connector (where N=Hpatch×WpatchN = H_{\text{patch}} \times W_{\text{patch}}), a 2D bilinear interpolation function is applied along the spatial grid dimensions with a downsampling factor of 2 in each spatial dimension (an area reduction factor of 4). For a 336px×336px336\text{px} \times 336\text{px} image with patch size 1414, the visual token count decreases from N=576N = 576 (24×2424 \times 24) to N′=144N' = 144 (12×1212 \times 12). The compressed visual embeddings ev′∈R144×Dte'_v \in \mathbb{R}^{144 \times D_t} are concatenated with the prompt text tokens before LLM ingestion, yielding an approximate 3×3\times acceleration in inference speed and reducing stage-two instruction fine-tuning wall-clock time from 9 hours to 6.5 hours on 8 NVIDIA A100 GPUs.

  4. Knowl 4 — Dense Connector Multi-Stage Training Protocol

    experimental setup

    Training MLLMs equipped with the Dense Connector follows a two-stage pipeline:

    1. Stage 1 (Feature Alignment Pre-training):

      • Parameters updated: Only the randomly initialized Dense Connector (2-layer MLP); the pre-trained visual encoder and LLM are kept frozen.
      • Epochs: 1 epoch.
      • Global batch size: 256.
      • Learning rate: 1×10−31 \times 10^{-3}.
      • Datasets: Pre-training data from LLaVA-1.5 (558K image captions) or Mini-Gemini (1.2M image-text caption pairs).
    2. Stage 2 (Visual Instruction Fine-Tuning):

      • Parameters updated: Dense Connector and the LLM parameters; the visual encoder remains frozen.
      • Epochs: 1 epoch.
      • Global batch size: 128.
      • Learning rate: 2×10−52 \times 10^{-5}.
      • Datasets: Instruction tuning data from LLaVA-1.5 (665K conversations) or Mini-Gemini (1.5M conversations).
      • LoRA configuration: When scaling to 34B or 70B parameter models (e.g., Hermes-2-Yi-34B, Llama-3-70B-Instruct), LoRA fine-tuning is used with LoRA rank r=128r = 128 and LoRA scaling factor αLoRA=256\alpha_{\text{LoRA}} = 256.

    Hardware allocations: Models up to 13B parameters are trained on 8 NVIDIA A100 (40GB VRAM) GPUs; 34B and 70B models are trained on 32 NVIDIA A100 (80GB VRAM) GPUs.

  5. Knowl 5 — Ablation of Dense Connector Layer Selection and Integration Variants

    data/table

    Ablation experiments evaluated on a 24-layer CLIP-ViT-L-336px visual encoder with a Vicuna-7B LLM backbone demonstrate the comparative performance of Sparse Token Integration (STI), Sparse Channel Integration (SCI), and Dense Channel Integration (DCI) against the single-layer baseline (LLaVA-1.5 using layer 24 only):

    Model Layer Index GQA VQAv2\text{VQA}^{\text{v2}} SQAI\text{SQA}^{\text{I}} VQAT\text{VQA}^{\text{T}} POPE MMB MMV LBW
    Baseline 24 62.0 78.5 66.8 58.2 85.9 64.3 31.1 65.4
    + STI 8, 16, 24 63.3 79.1 68.0 58.0 85.8 67.2 30.9 65.5
    + STI 8, 16, 20, 24 63.0 79.1 68.0 58.8 85.9 67.6 30.8 65.7
    + SCI 8, 16, 24 63.7 79.2 68.9 58.2 86.1 66.2 32.2 66.0
    + SCI 16, 24 63.0 79.0 67.6 58.2 86.0 65.6 31.7 65.6
    + SCI 8, 16, 20, 24 63.6 79.2 67.0 58.1 86.0 65.8 31.9 66.0
    + DCI (1-8),(9-16),(17-24) 63.6 79.3 67.8 58.6 86.3 66.5 32.6 66.0
    + DCI (1-12),(13-24) 63.8 79.5 69.5 59.2 86.6 66.8 32.7 66.1

    All multi-layer integration schemes surpass the single-layer baseline across benchmarks. DCI with 2 groups (1–12 and 13–24) attains the highest gains across GQA (+1.8%), VQAv2\text{VQA}^{\text{v2}} (+1.0%), SQAI\text{SQA}^{\text{I}} (+2.7%), VQAT\text{VQA}^{\text{T}} (+1.0%), POPE (+0.7%), MMB (+2.5%), MM-Vet (+1.6%), and LLaVA-Bench-in-the-Wild (+0.7%).

  6. Knowl 6 — Scalability and Compatibility Across Visual Encoders, Data Scales, and Resolutions

    data/table

    The Dense Connector scales across different visual backbones (CLIP-ViT-L vs. SigLIP-ViT-SO), training data volumes (LLaVA-1.5 0.5M+0.6M vs. Mini-Gemini 1.2M+1.5M), dual-encoder setups, and dynamic high-resolution architectures (AnyRes):

    Method VE Res. PT+IT LLM GQA SQAI\text{SQA}^{\text{I}} VQAT\text{VQA}^{\text{T}} MMB
    LLaVA CLIP-L 336 0.5M+0.6M Vicuna-7B 62.0 66.8 58.2 64.3
    LLaVA CLIP-L 336 0.5M+0.6M Vicuna-13B 63.3 71.6 61.3 67.6
    DC (w/ LLaVA) CLIP-L 336 0.5M+0.6M Vicuna-7B 63.8 69.5 59.2 66.8
    DC (w/ LLaVA) SigLIP-SO 384 0.5M+0.6M Vicuna-7B 64.2 70.5 62.6 68.4
    DC (w/ LLaVA) SigLIP-SO 384 0.5M+0.6M Vicuna-13B 65.4 73.0 64.7 71.4
    DC (w/ LLaVA) SigLIP-SO 384 1.2M+1.5M Vicuna-7B 63.8 72.9 64.6 71.7
    DC (w/ LLaVA) SigLIP-SO 384 1.2M+1.5M Vicuna-13B 64.6 77.1 65.0 74.4
    MGM CLIP-L+ConvX-L 336+768 1.2M+1.5M Vicuna-7B 62.6 70.4 65.2 69.3
    MGM CLIP-L+ConvX-L 336+768 1.2M+1.5M Vicuna-13B 63.4 72.6 65.9 68.5
    DC (w/ MGM) CLIP-L+ConvX-L 336+768 1.2M+1.5M Vicuna-7B 63.3 70.7 66.0 70.7
    DC (w/ MGM) CLIP-L+ConvX-L 336+768 1.2M+1.5M Vicuna-13B 64.2 74.9 66.7 70.7
    LLaVA-NeXT CLIP-L AnyRes 0.5M+0.6M Vicuna-7B 64.0 69.5 64.5 66.5
    DC (w/ LLaVA) CLIP-L AnyRes 0.5M+0.6M Vicuna-7B 64.6 70.5 65.6 67.4
    DC (w/ LLaVA) SigLIP-SO AnyRes 0.5M+0.6M Vicuna-7B 64.8 69.3 66.5 67.2

    Key findings include:

    1. Upgrading from CLIP-L to SigLIP-SO with Dense Connector improves MMB from 66.8% to 68.4% (7B) and 71.4% (13B).
    2. Increasing data to 1.2M+1.5M boosts Vicuna-13B performance on SQAI\text{SQA}^{\text{I}} to 77.1% and MMB to 74.4%.
    3. Incorporating DC into Mini-Gemini (applying DCI to the CLIP branch while retaining high-resolution ConvNeXt features) outperforms standard MGM across all benchmarks.
    4. Extending DC to LLaVA-NeXT with AnyRes dynamic high-resolution outperforms the base LLaVA-NeXT model across GQA, SQAI\text{SQA}^{\text{I}}, VQAT\text{VQA}^{\text{T}}, and MMB.
  7. Knowl 7 — State-of-the-Art Evaluation Across LLM Scales (2.7B to 70B)

    data/table

    Evaluation of Dense Connector models scaled from 2.7B to 70B parameters across 8 multimodal benchmarks shows consistent improvements over baseline architectures:

    Model LLM SQAI\text{SQA}^{\text{I}} MMB MMEP\text{MME}^{\text{P}} MM-Vet MMMUv\text{MMMU}^{\text{v}} Math LLaVAW\text{LLaVA}^{\text{W}} GQA
    TinyLLaVA Phi2-2.7B 69.9 – – 32.1 – – 67.9 61.3
    LLaVA-v1.5 Vicuna-13B 71.6 67.7 1531 36.1 36.4 27.6 72.5 63.3
    Mini-Gemini Vicuna-13B 72.6 68.5 1565 46.0 38.1 37.0 87.7 63.4
    LLaVA-NeXT Vicuna-13B 73.6 70.0 1575 48.4 36.2 35.3 87.3 65.4
    DC (0.5M+0.6M) Phi2-2.7B 70.3 70.5 1487 33.8 36.6 28.2 65.1 61.5
    DC (0.5M+0.6M) Vicuna-7B 70.5 68.4 1523 35.4 36.7 25.5 67.4 64.4
    DC (0.5M+0.6M) Vicuna-13B 73.0 71.4 1569 41.6 34.3 29.6 73.6 65.4
    DC (0.5M+0.6M) Llama3-8B 75.2 74.4 1558 34.6 40.4 28.6 68.8 65.1
    DC (0.5M+0.6M) Yi-34B 80.5 77.7 1588 41.0 47.1 33.5 75.1 63.9
    DC (0.5M+0.6M) Llama3-70B 82.4 79.4 1622 46.1 47.0 32.9 74.5 64.0
    DC (1.2M+1.5M) Vicuna-13B 77.1 74.4 1579 47.8 37.2 36.5 88.9 64.6
    DC (AnyRes) Yi-34B 78.0 81.2 1696 59.2 51.8 40.0 97.7 66.6

    DC with Phi2-2.7B outperforms TinyLLaVA (70.5 vs. unlisted MMB, 33.8 vs. 32.1 on MM-Vet). DC with Vicuna-13B under standard data matches or exceeds Mini-Gemini and LLaVA-v1.5 across GQA, SQAI\text{SQA}^{\text{I}}, and MMB. When scaling to Llama3-70B, the model achieves 82.4% on SQAI\text{SQA}^{\text{I}} and 79.4% on MMB. Combined with dynamic high resolution (AnyRes) and 1.2M+1.5M training data on Yi-34B, DC achieves 81.2% on MMB, 59.2% on MM-Vet, 51.8% on MMMUv\text{MMMU}^{\text{v}}, 40.0% on MathVista, and 97.7% on LLaVAW\text{LLaVA}^{\text{W}}.

  8. Knowl 8 — Token Efficiency Comparison Between Efficient Dense Connector and Existing Compressors

    data/table

    Comparison of the Efficient Dense Connector against LLaVA-1.5, Qwen-VL-Chat, and TokenPacker on image understanding benchmarks using a Vicuna-7B backbone:

    Method Res. #Token GQA VQAv2\text{VQA}^{\text{v2}} SQAI\text{SQA}^{\text{I}} VQAT\text{VQA}^{\text{T}} MMB MMV
    LLaVA-v1.5 336 576 62.0 78.5 66.8 58.2 64.3 31.1
    Qwen-VL-Chat 448 256 57.5 68.2 61.5 – – –
    TokenPacker 336 144 61.9 77.9 – – 65.1 33.0
    Dense Connector 336 144 62.8 79.4 68.8 58.1 67.6 34.4

    While reducing visual token count by 75%75\% (from 576 to 144 tokens), Dense Connector maintains or exceeds the accuracy of the 576-token LLaVA-1.5 baseline across GQA (62.8 vs. 62.0), VQAv2\text{VQA}^{\text{v2}} (79.4 vs. 78.5), SQAI\text{SQA}^{\text{I}} (68.8 vs. 66.8), MMB (67.6 vs. 64.3), and MM-Vet (34.4 vs. 31.1). At the identical budget of 144 tokens, Dense Connector surpasses TokenPacker by +0.9% on GQA, +1.5% on VQAv2\text{VQA}^{\text{v2}}, +2.5% on MMB, and +1.4% on MM-Vet.

  9. Knowl 9 — Zero-Shot Video Question Answering via Training-Free Extension

    empirical result

    By applying the FreeVA training-free frame sampling strategy, image-trained Dense Connector models are evaluated on open-ended video QA benchmarks without any video-specific training. A video is represented by uniformly sampling TT frames, encoding each frame through the frozen visual encoder and Dense Connector into embedding sequence {ev1,…,evT}\{e_{v_1}, \dots, e_{v_T}\}, and concatenating them as input to the LLM.

    Under evaluation using the GPT-3.5-Turbo-0125 (JAN) evaluator:

    1. Vicuna-7B (DC+FreeVA): Achieves 75.0% accuracy (score 4.1) on MSVD-QA, 58.4% accuracy (score 3.5) on MSRVTT-QA, and 52.2% accuracy (score 3.5) on ActivityNet-QA. On the Video-ChatGPT benchmark, it scores 2.80 on Correctness of Information (CI), 2.51 on Detail Orientation (DO), 3.17 on Contextual Understanding (CU), 2.22 on Temporal Understanding (TU), and 3.05 on Consistency (CO).
    2. Vicuna-13B (DC+FreeVA): Achieves 75.1% accuracy (score 4.1) on MSVD-QA, 60.8% accuracy (score 3.5) on MSRVTT-QA, and 52.6% accuracy (score 3.5) on ActivityNet-QA, outperforming the LLaVA-1.5 + FreeVA 13B baseline on MSVD-QA (74.4%) and ActivityNet-QA (51.6%).
    3. Yi-34B (DC+FreeVA): Attains state-of-the-art zero-shot performance across benchmarks with 77.4% accuracy (score 4.2) on MSVD-QA, 62.1% accuracy (score 3.6) on MSRVTT-QA, 55.8% accuracy (score 3.6) on ActivityNet-QA, and Video-ChatGPT metrics of CI 3.00, DO 2.53, CU 3.25, TU 2.65, and CO 2.92.
  10. Knowl 10 — Impact of Parameterized Multi-Layer Fusion and Vision Transformer Fine-Tuning

    empirical result

    Ablation experiments exploring additional parameterized fusion mechanisms and visual encoder fine-tuning indicate:

    1. Parameterized vs. Non-Parameterized Connector Layers: Adding learnable operations into multi-layer integration—such as 1D convolutional downsampling for STI (ev=MLP(Concat([Conv1D(Vl1),…,VL]))e_v = \text{MLP}(\text{Concat}([\text{Conv1D}(V_{l_1}), \dots, V_L]))), 2D convolutions prior to concatenation for SCI (ev=MLP(Concat([Conv2D(Vl1),…,Conv2D(VL)]))e_v = \text{MLP}(\text{Concat}([\text{Conv2D}(V_{l_1}), \dots, \text{Conv2D}(V_L)]))), or shared linear projection layers with LayerNorm for DCI (GVg=1M∑Linear(Ln(Vi))GV_g = \frac{1}{M}\sum \text{Linear}(\text{Ln}(V_i)))—does not improve performance over parameter-free averaging/summation. For example, on GQA and MMBench with a CLIP-L backbone on Vicuna-7B:

      • STI w/ 1D Conv scores 62.6% (GQA) and 65.0% (MMB) vs. 63.3% and 67.2% for STI with average pooling.
      • SCI w/ 2D Conv scores 63.6% (GQA) and 66.0% (MMB) vs. 63.7% and 66.2% for SCI without convolution.
      • DCI w/ Linear scores 63.7% (GQA) and 66.4% (MMB) vs. 63.8% and 66.8% for standard DCI. Randomly initialized parameterized fusion layers hinder optimization convergence when training on limited alignment data.
    2. Fine-Tuning the Visual Encoder (ViT): Fine-tuning the ViT during stage-two instruction tuning with a reduced learning rate (2×10−62 \times 10^{-6}) improves performance on visually intensive benchmarks (CLIP-L + Vicuna-7B on MMBench increases from 66.8% to 68.6%, and TextVQA increases from 59.2% to 60.2%), but causes minor degradation on visually less dependent reasoning benchmarks (ScienceQA drops from 69.5% to 67.4%).

  11. Knowl 11 — Limitations of Dense Connector Design and Zero-Shot Video Adaptation

    limitation

    The authors identify two specific limitations of the proposed approach:

    1. Inability to Benefit from Learnable Fusion Modules: The current Dense Connector instantiations (STI, SCI, DCI) rely entirely on non-parameterized feature grouping and pooling before the final 2-layer MLP projection. Introducing parameterized modules (e.g., convolutional downsampling or linear projection layers) fails to enhance downstream performance due to optimization convergence difficulties during the initial alignment phase.
    2. Linguistic Modality Confusion in Training-Free Video Inference: When adapting image-trained Dense Connector models directly to video dialogues without video-specific fine-tuning, the LLM occasionally exhibits modality confusion in its textual responses, referring to input video sequences as 'The image' rather than 'The video'.

Coverage note — Qualitative visual dialogue examples from Figures 3 and 5-13 were omitted as they represent illustrative demonstrations of model behavior rather than independent technical contributions.

References

  1. 1.OpenAI. Chatgpt. https://openai.com/blog/chatgpt/, 2023.
  2. 2.OpenAI. Gpt-4v(ision) system card. 2023.
  3. 3.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  4. 4.Wenhao Wu, Huanjin Yao, Mengxi Zhang, Yuxin Song, Wanli Ouyang, and Jingdong Wang. Gpt4vis: What can gpt-4 do for zero-shot visual recognition? arXiv preprint arXiv:2311.15732, 2023.
  5. 5.Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9(1):1, 2023.
  6. 6.Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, et al. Can gpt-4v (ision) serve medical applications? case studies on gpt-4v for multimodal medical diagnosis. arXiv preprint arXiv:2310.09909, 2023.
  7. 7.Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi. Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration. arXiv preprint arXiv:2311.12015, 2023.
  8. 8.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  9. 9.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  10. 10.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  11. 11.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  12. 12.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6, 2023.
  13. 13.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022.
  14. 14.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023.
  15. 15.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  16. 16.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023.
  17. 17.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023.
  18. 18.Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814, 2024.
  19. 19.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024.
  20. 20.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  21. 21.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  22. 22.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  23. 23.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  24. 24.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  25. 25.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.
  26. 26.Bo Li, Peiyuan Zhang, Jingkang Yang, Yuanhan Zhang, Fanyi Pu, and Ziwei Liu. Otterhd: A high-resolution multi-modality model. arXiv preprint arXiv:2311.04219, 2023.
  27. 27.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023.
  28. 28.Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al. Mm1: Methods, analysis & insights from multimodal llm pre-training. arXiv preprint arXiv:2403.09611, 2024.
  29. 29.Dongsheng Jiang, Yuchen Liu, Songlin Liu, Xiaopeng Zhang, Jin Li, Hongkai Xiong, and Qi Tian. From clip to dino: Visual encoders shout in multi-modal large language models. arXiv preprint arXiv:2310.08825, 2023.
  30. 30.Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms. arXiv preprint arXiv:2401.06209, 2024.
  31. 31.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023.
  32. 32.Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2818–2829, 2023.
  33. 33.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. Filip: Fine-grained interactive language-image pre-training. In ICLR, 2021.
  34. 34.Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Yuxin Song, Weiping Wang, Xiangbo Shu, Xiangyang Ji, and Jingdong Wang. Uatvr: Uncertainty-adaptive text-video retrieval. In ICCV, 2023.
  35. 35.Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang. Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10704–10713, 2023.
  36. 36.Wenhao Wu, Zhun Sun, Yuxin Song, Jingdong Wang, and Wanli Ouyang. Transferring vision-language models for visual recognition: A classifier perspective. International Journal of Computer Vision, pages 1–18, 2023.
  37. 37.Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, and Wanli Ouyang. Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models. In CVPR, pages 6620–6630, 2023.
  38. 38.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
  39. 39.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020.
  41. 41.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  43. 43.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  44. 44.Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al. Phi-2: The surprising power of small language models. Microsoft Research Blog, 2023.
  45. 45.Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024.
  46. 46.Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al. Stable lm 2 1.6 b technical report. arXiv preprint arXiv:2402.17834, 2024.
  47. 47.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  48. 48.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  49. 49.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  50. 50.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems, 36, 2024.
  51. 51.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023.
  52. 52.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023.
  53. 53.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023.
  54. 54.Wenhao Wu. Freeva: Offline mllm as training-free video assistant. arXiv preprint arXiv:2405.07798, 2024.
  55. 55.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  56. 56.Wentong Li, Yuqian Yuan, Jian Liu, Dongqi Tang, Song Wang, Jianke Zhu, and Lei Zhang. Tokenpacker: Efficient visual projector for multimodal llm. arXiv preprint arXiv:2407.02392, 2024.
  57. 57.Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:2402.03766, 2024.
  58. 58.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709, 2019.
  59. 59.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017.
  60. 60.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022.
  61. 61.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019.
  62. 62.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023.
  63. 63.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.
  64. 64.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023.
  65. 65.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  66. 66.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023.
  67. 67.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  68. 68.David Chen and William B Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011.
  69. 69.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, 2015.
  70. 70.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016.
  71. 71.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022.
  72. 72.Baichuan Zhou, Ying Hu, Xi Weng, Junlong Jia, Jie Luo, Xien Liu, Ji Wu, and Lei Huang. Tinyllava: A framework of small-scale large multimodal models. arXiv preprint arXiv:2402.14289, 2024.
  73. 73.Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. arXiv preprint arXiv:2311.04257, 2023.
  74. 74.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. arXiv preprint arXiv:2311.17043, 2023.
  75. 75.Peng Gao, Renrui Zhang, Chris Liu, Longtian Qiu, Siyuan Huang, Weifeng Lin, Shitian Zhao, Shijie Geng, Ziyi Lin, Peng Jin, et al. Sphinx-x: Scaling data and parameters for a family of multi-modal large language models. arXiv preprint arXiv:2402.05935, 2024.
  76. 76.XTuner Contributors. Xtuner: A toolkit for efficiently fine-tuning llm. https://github.com/InternLM/xtuner, 2023.
  77. 77.Jiachen Li, Xinyao Wang, Sijie Zhu, Chia-Wen Kuo, Lu Xu, Fan Chen, Jitesh Jain, Humphrey Shi, and Longyin Wen. Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts. arXiv preprint arXiv:2405.05949, 2024.
  78. 78.Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. arXiv preprint arXiv:2312.07533, 2023.
  79. 79.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. Advances in Neural Information Processing Systems, 35:124–141, 2022.
  80. 80.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023.
  81. 81.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  82. 82.Yizhou Wang, Ruiyi Zhang, Haoliang Wang, Uttaran Bhattacharya, Yun Fu, and Gang Wu. Vaquita: Enhancing alignment in llm-assisted video understanding. arXiv preprint arXiv:2312.02310, 2023.
  83. 83.Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan, Thomas H Li, and Ge Li. One for all: Video conversation is feasible without video instruction tuning. arXiv preprint arXiv:2309.15785, 2023.
  84. 84.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  85. 85.Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024.
  86. 86.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.

Citation

MLA
Yao, H., et al. “Dense Connector for MLLMs”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 33108–40, https://proceedings.neurips.cc/paper_files/paper/2024/file/3a10c46572628d58cb44fb705f25cbbf-Paper-Conference.pdf.
APA
Yao, H., Wu, W., Yang, T., Song, Y., Zhang, M., Feng, H., Sun, Y., Li, Z., Ouyang, W., & Wang, J. (2024). Dense Connector for MLLMs. Advances in Neural Information Processing Systems, 37, 33108–33140. https://proceedings.neurips.cc/paper_files/paper/2024/file/3a10c46572628d58cb44fb705f25cbbf-Paper-Conference.pdf
Chicago
Yao, H., W. Wu, T. Yang, et al. 2024. “Dense Connector for MLLMs”. Advances in Neural Information Processing Systems 37: 33108–40. https://proceedings.neurips.cc/paper_files/paper/2024/file/3a10c46572628d58cb44fb705f25cbbf-Paper-Conference.pdf.
Harvard
Yao, H. et al. (2024) “Dense Connector for MLLMs”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 33108–33140. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/3a10c46572628d58cb44fb705f25cbbf-Paper-Conference.pdf.
Vancouver
1. Yao H, Wu W, Yang T, Song Y, Zhang M, Feng H, Sun Y, Li Z, Ouyang W, Wang J (2024) Dense Connector for MLLMs. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 33108–33140

BibTeX

@inproceedings{yao2024dense,
  title = {Dense Connector for MLLMs},
  author = {Yao, Huanjin and Wu, Wenhao and Yang, Taojiannan and Song, Yuxin and Zhang, Mengxi and Feng, Haocheng and Sun, Yifan and Li, Zhiheng and Ouyang, Wanli and Wang, Jingdong},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {33108-33140},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/3a10c46572628d58cb44fb705f25cbbf-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors