Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Bin LinBin ZhuYang YeMunan NingPeng JinLi Yuan

article2023EMNLP1,774 citations

Introduces Video-LLaVA, a vision-language model that aligns image and video features into a single representation before projecting them to a large language model, demonstrating that joint multimodal training outperforms specialized single-modality systems across major image and video benchmarks.

Listen

Most large artificial intelligence models are designed to handle either text and static images or text and video streams separately, leading to fragmented visual understanding. When separate systems process images and videos independently, the core language model struggles to connect insights across media types, hindering unified multi-modal reasoning.

The article demonstrates and evaluates Video-LLaVA, a unified vision-language framework designed to align image and video data before projecting them into a central language model. The main objective was to establish a single baseline model capable of mutual cross-modality learning across both visual domains without relying on separate, disjointed pipelines.

To accomplish this, the authors initialized image and video encoders using LanguageBind to map both visual media directly into a shared language feature space. They evaluated the model on four video reasoning benchmarks and nine standard image evaluation toolkits. The training process followed a two-stage approach: visual pretraining using roughly 558,000 image-text pairs and 702,000 video-text pairs, followed by instruction tuning across 665,000 image samples and 100,000 video samples.

The findings show that unifying visual representations prior to projection significantly improves multi-modal intelligence. First, Video-LLaVA surpassed Video-ChatGPT on major video reasoning benchmarks, increasing accuracy by 5.8% on MSVD, 9.9% on MSRVTT, 18.6% on TGIF, and 10.1% on ActivityNet. Second, the 7-billion parameter Video-LLaVA model outperformed larger models on broad image benchmarks, including exceeding the 80-billion parameter IDEFICS model by 6.4% on MMBench. Third, joint training on both images and videos reduced object hallucination and improved visual conversation compared to image-only and video-only training setups.

These results demonstrate that joint visual alignment provides a performance advantage over specialized single-modality systems. Adopting this unified architecture can lower engineering complexity and consolidate infrastructure costs by eliminating the need to deploy separate image and video comprehension models.

Organizations developing multi-modal AI systems should prioritize unified visual pre-alignment strategies over separate projection pathways. Future implementation efforts should explore token compression methods to improve processing efficiency and expand the architecture to temporal timestamps and non-visual sensor data like depth and infrared feeds.

A key limitation is that Video-LLaVA samples only eight uniform frames per video, which constrains its ability to capture fine-grained details in long videos. Additionally, model training required substantial computing resources, taking three to four days on eight high-end graphics processing units. Overall, confidence in the reported benchmark gains is high, though caution is advised when deploying the current architecture on extended video sequences.

Cover for Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Abstract

The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers. In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM. As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other. Video-LLaVA achieves superior performances on a broad range of 9 image benchmarks across 5 image question-answering datasets and 4 image benchmark toolkits. Additionally, our Video-LLaVA also outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively. Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos. We aim for this work to provide modest insights into the multi-modal inputs for the LLM. Code address: \href{this https URL}

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Language Models
  • 2.2 Large Vision-Language Models
  • 2.2.1 LLMs as scheduler
  • 2.2.2 LLMs as decoder
  • 3 Video-LLaVA
  • 3.1 Model Structure
  • 3.1.1 Framework Overview
  • 3.1.2 United Visual Representation
  • 3.1.3 Alignment Before Projection
  • 3.2 Training Pipeline
  • 3.2.1 Understanding Training
  • 3.2.2 Instruction Tuning
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.1.1 Data Details
  • 4.1.2 Model Settings
  • 4.1.3 Training Details
  • 4.2 Quantitative Evaluation
  • 4.2.1 Zero-shot Video Understanding
  • 4.2.2 Zero-shot Image Question-answering
  • 4.2.3 Object Hallucination Evaluation
  • 4.3 Ablation Results
  • 4.3.1 Alignment Before Projection
  • 4.3.2 For Video Understanding
  • 4.3.3 For Image Understanding
  • 4.3.4 Joint Training
  • 4.3.5 For Video Understanding
  • 4.3.6 For Image Understanding
  • 5 Limitation and Future Directions
  • 5.1 Limitation
  • 5.2 Future Directions
  • 6 Conclusion
  • References
  • A Example Appendix
  • A.1 Training Setting
  • A.2 Exhibition Board

Knowls

  1. Knowl 1 — Video-LLaVA Architecture and Visual Representation Pre-Alignment

    model/method

    Video-LLaVA is a large vision-language model (LVLM) architecture designed to process both static images and videos within a unified visual representation space before projecting them into a Large Language Model (LLM).

    The framework consists of four primary components:

    1. Unified Modality Encoders (fVf_V): Visual encoders for images and videos initialized from LanguageBind (based on OpenCLIP-L/14). In LanguageBind, images and language are aligned in a shared feature space, and video representations are aligned to the same textual space using 3 million video-text pairs from VIDAL-10M. This pre-aligns image and video feature spaces into an emergent unified visual space prior to any projection.
    2. Shared Visual Projection Layer (fPf_P): A 2-layer Multi-Layer Perceptron (MLP) with Gaussian Error Linear Unit (GeLU) activation functions that maps the unified visual features from both image and video encoders into the language embedding dimension.
    3. Word Embedding Layer (fTf_T): A standard text embedding layer mapping input text tokens to the LLM token embedding space (utilizing a 32,000-token vocabulary from LLaMA).
    4. Large Language Model (fLf_L): A foundational autoregressive LLM backend, specifically Vicuna-7B v1.5, which receives concatenated visual and textual token representations to generate text responses.
  2. Knowl 2 — Two-Stage Joint Training Pipeline of Video-LLaVA

    model/method

    Video-LLaVA is trained via a two-stage procedure where every training batch contains a randomized mixture of image and video samples:

    1. Stage 1: Understanding Pretraining: Focuses on learning basic visual concepts and aligning visual signals to the language embedding space. The training data consists of single-turn image/video-text pairs focused on concise captions (558k LAION-CC-SBU image-text pairs filtered by LLaVA and 702k WebVid video-text pairs from Valley). In this stage, the visual encoders and the LLM are kept frozen; only the shared projection layer is trained for 1 epoch.
    2. Stage 2: Instruction Tuning: Focuses on enabling multi-turn conversation, detailed captioning, and complex visual reasoning. The dataset comprises 665k multi-turn image-text instruction pairs from LLaVA 1.5 and 100k video-text instruction pairs from Video-ChatGPT. In this stage, both the shared visual projection layer and the LLM are updated end-to-end for 1 epoch, while the visual encoders remain frozen.
  3. Knowl 3 — Autoregressive Multi-Modal Likelihood Formulation and Multi-Turn Prompting

    equation

    Given a textual input XTX_T and visual input signals XVX_V (where XVX_V can be an image or an 8-frame video sequence), the feature embeddings are computed as:

    ZT=fT(XT),ZV=fP(fV(XV))Z_T = f_T(X_T), \quad Z_V = f_P(f_V(X_V))

    where fVf_V is the visual encoder, fPf_P is the shared projection layer, and fTf_T is the text embedding layer.

    The conditional generation probability of the response sequence XA={XA[1],XA[2],…,XA[L]}X_A = \{X_A^{[1]}, X_A^{[2]}, \dots, X_A^{[L]}\} of length LL parameterized by trainable parameters θ\theta is formulated as:

    p(XA∣XV,XT)=∏i=1Lpθ(XA[i]∣ZV,ZT,XA[1:i−1])p(X_A \mid X_V, X_T) = \prod_{i=1}^L p_\theta\left(X_A^{[i]} \mid Z_V, Z_T, X_A^{[1:i-1]}\right)

    For a multi-turn conversation with rr rounds where round rr has query XqrX_q^r and response XarX_a^r, the full text context XTrX_T^r provided to the model at round rr is constructed iteratively via:

    XTr={Xq1,r=1Concat(XTr−1,XAr−1,Xqr),r>1X_T^r = \begin{cases} X_q^1, & r = 1 \\ \text{Concat}\left(X_T^{r-1}, X_A^{r-1}, X_q^r\right), & r > 1 \end{cases}

  4. Knowl 4 — Zero-Shot Video Question Answering Performance

    data/table

    Video-LLaVA was evaluated in a zero-shot setting across four open-ended video question-answering benchmarks using the Video-ChatGPT evaluation pipeline, where accuracy (%) and qualitative score (on a 1–5 scale) are assessed by a GPT-3.5-turbo assistant.

    Methods LLM size MSVD-QA MSRVTT-QA TGIF-QA ActivityNet-QA
    Acc Score Acc Score Acc Score Acc Score
    FrozenBiLM 1B 32.2 - 16.8 - 41.0 - 24.7 -
    VideoChat 7B 56.3 2.8 45.0 2.5 34.4 2.3 - 2.2
    LLaMA-Adapter 7B 54.9 3.1 43.8 2.7 - - 34.2 2.7
    Video-LLaMA 7B 51.6 2.5 29.6 1.8 - - 12.4 1.1
    Video-ChatGPT 7B 64.9 3.3 49.3 2.8 51.4 3.0 35.2 2.7
    Chat-UniVi 7B 65.0 3.6 54.6 3.1 60.3 3.4 45.8 3.2
    Video-LLaVA 7B 70.7 3.9 59.2 3.5 70.0 4.0 45.3 3.3

    Video-LLaVA-7B outperforms the dedicated video model Video-ChatGPT by 5.8% on MSVD-QA, 9.9% on MSRVTT-QA, 18.6% on TGIF-QA, and 10.1% on ActivityNet-QA.

  5. Knowl 5 — Image Understanding and Instruction Tuning Benchmark Results

    data/table

    Video-LLaVA (7B, input image resolution 224×224224 \times 224) was evaluated on five standard visual question answering benchmarks (VQA-v2, GQA, VisWiz, ScienceQA-IMG, TextVQA) and four instruction-tuning benchmark toolkits (POPE, MMBench, LLaVA-Bench In-the-Wild, MM-Vet).

    Methods LLM Res. VQAv2^{v2} GQA VisWiz SQAI^I VQAT^T POPE MMB LLaVAW^W MM-Vet
    BLIP-2 V-13B 224 41.0 41.0 19.6 61.0 42.5 85.3 - 38.1 22.4
    InstructBLIP V-13B 224 - 49.5 33.4 63.1 50.7 78.9 - 58.2 25.6
    IDEFICS-80B L-65B 224 60.0 45.2 36.0 - 30.9 - 54.5 - -
    MiniGPT-4 L-7B 224 - 30.8 47.5 25.4 19.4 - 23.0 - 22.1
    IDEFICS-9B L-7B 224 50.9 38.4 35.5 - 25.9 - 48.2 - -
    mPLUG-Owl L-7B 224 - 14.0 39.0 2.8 38.8 - 46.6 - -
    Otter L-7B 224 - 38.1 50.0 27.2 21.2 - 32.6 - 24.6
    InstructBLIP V-7B 224 - 49.2 34.5 60.5 50.1 - 36.0 60.9 26.2
    LLaVA-1.5†^{\dagger} V-7B 224 72.3 56.9 47.8 67.9 49.2 83.3 59.5 63.3 25.7
    Video-LLaVA V-7B 224 74.7 60.3 48.1 66.4 51.8 84.4 60.9 73.1 32.0

    Note: L denotes LLaMA, V denotes Vicuna. †\dagger denotes LLaVA-1.5 reproduced at 224×224224 \times 224 using the LanguageBind-Image encoder. Despite being trained simultaneously on video and image data, Video-LLaVA outperforms image-specific baselines (such as InstructBLIP-7B and LLaVA-1.5†^{\dagger}) across nearly all benchmarks.

  6. Knowl 6 — Object Hallucination Evaluation on POPE

    data/table

    Zero-shot object hallucination was assessed using the Polling-based Object Probing Evaluation (POPE) pipeline across three question distribution settings: Adversarial, Popular, and Random. Metrics include Accuracy (%), F1-Score (%), and proportion of positive ("Yes") responses.

    Methods LLM Adversarial Popular Random
    Acc F1 Yes Acc F1 Yes Acc F1 Yes
    MiniGPT-4 V-13B 66.6 71.4 66.7 68.3 72.2 64.1 77.8 78.9 54.8
    InstructBLIP V-13B 74.4 78.5 69.0 81.4 83.5 62.6 88.7 89.3 55.2
    MM-GPT L-7B 50.0 66.7 100.0 50.0 66.7 100.0 50.0 66.7 100.0
    mPLUG-Owl L-7B 50.7 66.8 98.7 50.9 66.9 98.6 54.0 66.4 95.6
    Chat-UniVi V-7B 55.6 68.7 91.6 56.4 69.0 90.8 73.9 79.3 74.6
    LLaVA-1.5†^{\dagger} L-7B 84.3 83.2 43.5 79.8 79.4 48.0 85.7 84.8 43.0
    Video-LLaVA V-7B 81.6 80.8 45.8 85.3 84.0 42.1 86.2 85.2 42.0

    Note: †\dagger indicates reproduction of LLaVA-1.5 with the LanguageBind-Image encoder. Video-LLaVA significantly outperforms previous 7B-scale models (MM-GPT, mPLUG-Owl, Chat-UniVi) and larger 13B models (MiniGPT-4), achieving balanced "Yes" ratios closer to 50% and higher F1 scores on Popular and Random subsets.

  7. Knowl 7 — Ablation Study on Alignment Before Projection

    empirical result

    To evaluate the importance of pre-aligning visual spaces before the projection layer, Video-LLaVA's unified visual representation (LanguageBind) was compared against architectures with separated visual representations where the image encoder is replaced while keeping the LanguageBind video encoder fixed:

    1. Separated-MAE: Uses a Masked Autoencoder (MAE) image feature extractor (not trained on multi-modal data).
    2. Separated-CLIP: Uses standard CLIP-L/14 (multimodal, but not pre-aligned with the LanguageBind video encoder).
    3. United (Video-LLaVA): Uses LanguageBind image and video encoders pre-aligned into the same space.
    Methods VQAv2^{v2} GQA VisWiz SQAI^I VQAT^T POPE MMB LLaVAW^W MM-Vet
    Separated-MAE 66.0 55.4 42.5 65.0 44.2 80.8 45.7 35.9 20.0
    Separated-CLIP 74.6 59.9 47.8 67.3 51.5 84.4 60.2 68.9 30.6
    United (Video-LLaVA) 74.7 60.3 48.1 66.4 51.8 84.4 60.9 73.1 32.0

    The United setup consistently outperforms Separated-CLIP and Separated-MAE across both image tasks (gaining +4.2% on LLaVA-Bench and +1.4% on MM-Vet over Separated-CLIP) and video QA benchmarks (improving accuracy and GPT score across MSVD, TGIF, MSRVTT, and ActivityNet), demonstrating that pre-aligning modalities into a shared visual space reduces projection difficulty for the LLM.

  8. Knowl 8 — Ablation Study on Joint Training Synergy of Images and Videos

    empirical result

    Experiments isolating video and image data during training demonstrate that joint training produces mutual improvements across modalities:

    1. Video Understanding Gains from Image Data: Training Video-LLaVA on video data only (designated Video-LLaVA∗\text{Video-LLaVA}^*) versus joint image-video training shows marked improvements across all video benchmarks:
    Methods MSVD-QA MSRVTT-QA TGIF-QA ActivityNet-QA
    Acc Score Acc Score Acc Score Acc Score
    Video-LLaVA∗\text{Video-LLaVA}^* (Video Only) 64.8 3.2 58.3 3.4 67.8 3.4 40.7 2.0
    Joint with Image 70.7 3.9 59.2 3.5 70.0 4.0 45.3 3.3
    Δ\Delta +5.9 +0.7 +0.9 +0.1 +2.2 +0.6 +4.6 +1.3
    1. Image Understanding Gains from Video Data: Comparing Video-LLaVA with an identically configured image-only model (LLaVA-1.5†^{\dagger} reproduced with LanguageBind-Image encoder at 224×224224 \times 224) shows improvements in 8 out of 9 image benchmarks, notably increasing LLaVA-Bench from 63.3 to 73.1, MM-Vet from 25.7 to 32.0, and POPE from 83.3 to 84.4.
  9. Knowl 9 — Experimental Training Configuration and Hyperparameters of Video-LLaVA

    experimental setup

    Video-LLaVA uses the following hyperparameter and architectural configurations:

    • Input Dimensions: Images cropped/resized to 224×224224 \times 224; videos processed by uniformly sampling 8 frames per video at 224×224224 \times 224 per frame.
    • Visual Feature Layer: Features extracted from layer -2 of the visual encoder.
    • Encoders: LanguageBind-Image and LanguageBind-Video-LoRA encoders (both frozen during Stage 1 and Stage 2).
    • Optimizer: AdamW with cosine learning rate schedule and 0.0 weight decay.
    • Distributed Framework: DeepSpeed ZeRO-2.
    • Warmup Ratio: 0.03 for both stages.
    • Stage 1 (Pretraining): 1 epoch, batch size 256, initial learning rate 1×10−31 \times 10^{-3}.
    • Stage 2 (Instruction Tuning): 1 epoch, batch size 128, initial learning rate 2×10−52 \times 10^{-5}.
  10. Knowl 10 — Limitations in Long Video Understanding and Computational Requirements

    limitation

    Video-LLaVA possesses two primary limitations identified by the authors:

    1. Long Video Temporal Resolution: Because the model uniformly samples only 8 frames per video, it suffers information loss on long-form video content. For example, on ActivityNet-QA, models utilizing dynamic frame/token sampling (e.g., Chat-UniVi at 45.8% accuracy) outperform Video-LLaVA (45.3% accuracy).
    2. Computational Expense: Full training of the two stages requires substantial compute, taking approximately 3 to 4 days on 8 NVIDIA A100 (80GB) GPUs.

Coverage note — None was omitted; all key architectural concepts, mathematical formulations, empirical benchmark results, ablation studies, training settings, and explicit limitations have been captured.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738.
  4. 4.Bin Bi, Chenliang Li, Chen Wu, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2020. Palm: Pre-training an autoencoding&autoregressive language model for context-conditioned generation. arXiv preprint arXiv:2004.07159.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.David Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 190–200.
  7. 7.Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, and Bo Xu. 2023. X-llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages. arXiv preprint arXiv:2305.04160.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna.lmsys.org (accessed 14 April 2023).
  9. 9.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. Preprint, arXiv:2305.06500.
  10. 10.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. 2023. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394.
  11. 11.Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. 2023. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010.
  12. 12.Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15180–15190.
  13. 13.Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. 2023. Multimodal-gpt: A vision and language model for dialogue with humans. arXiv preprint arXiv:2305.04790.
  14. 14.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913.
  15. 15.Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. 2018. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617.
  16. 16.Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, et al. 2023. Imagebind-llm: Multi-modality instruction tuning. arXiv preprint arXiv:2309.03905.
  17. 17.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009.
  18. 18.Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415.
  19. 19.Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709.
  20. 20.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. 2021. Openclip. If you use this software, please cite it as below.
  21. 21.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017. Tgif-qa: Toward spatiotemporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2758–2766.
  22. 22.Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan. 2023. Chat-univi: Unified visual representation empowers large language models with image and video understanding. arXiv preprint arXiv:2311.08046.
  23. 23.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR.
  24. 24.Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh. 2023. Obelics: An open web-scale filtered dataset of interleaved image-text documents. Preprint, arXiv:2306.16527.
  25. 25.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023a. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726.
  26. 26.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pretraining with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  27. 27.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pretraining for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR.
  28. 28.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34:9694–9705.
  29. 29.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023c. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355.
  30. 30.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023d. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355.
  31. 31.Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Junwu Zhang, Munan Ning, and Li Yuan. 2024. Moe-llava: Mixture of experts for large vision-language models. arXiv preprint arXiv:2401.15947.
  32. 32.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744.
  33. 33.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  34. 34.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2023c. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281.
  35. 35.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521.
  36. 36.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Minghui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207.
  37. 37.Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu. 2023. Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration. arXiv preprint arXiv:2306.09093.
  38. 38.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424.
  39. 39.OpenAI. 2023. Gpt-4 technical report. Preprint, arXiv:2303.08774.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  41. 41.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  42. 42.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565.
  43. 43.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  44. 44.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326.
  45. 45.Dídac Surís, Sachit Menon, and Carl Vondrick. 2023. Vipergpt: Visual inference via python execution for reasoning. arXiv preprint arXiv:2303.08128.
  46. 46.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: An instruction-following llama model.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  49. 49.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671.
  50. 50.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296.
  51. 51.Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. 2023. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381.
  52. 52.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178.
  53. 53.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549.
  54. 54.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490.
  55. 55.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9127–9134.
  56. 56.Hang Zhang, Xin Li, and Lidong Bing. 2023a. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
  57. 57.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2023b. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199.
  58. 58.Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. 2023a. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:2310.01852.
  59. 59.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023b. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Lin, B., et al. “Video-LLaVA: Learning United Visual Representation by Alignment Before Projection”. arXiv, 2023, http://arxiv.org/abs/2311.10122v3.
APA
Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., & Yuan, L. (2023). Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. arXiv. http://arxiv.org/abs/2311.10122v3
Chicago
Lin, B., Y. Ye, B. Zhu, et al. 2023. “Video-LLaVA: Learning United Visual Representation by Alignment Before Projection”. arXiv. http://arxiv.org/abs/2311.10122v3.
Harvard
Lin, B. et al. (2023) “Video-LLaVA: Learning United Visual Representation by Alignment Before Projection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.10122v3.
Vancouver
1. Lin B, Ye Y, Zhu B, Cui J, Ning M, Jin P, Yuan L (2023) Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. arXiv

BibTeX

@article{lin2023video,
  title = {Video-LLaVA: Learning United Visual Representation by Alignment Before Projection},
  author = {Lin, Bin and Ye, Yang and Zhu, Bin and Cui, Jiaxi and Ning, Munan and Jin, Peng and Yuan, Li},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.10122v3},
  eprint = {2311.10122}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/