Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Hang ZhangXin LiLidong Bing

article2023EMNLP1,979 citations

Develops Video-LLaMA, an instruction-tuned multimodal framework that bridges frozen large language models with dedicated video and audio query transformers to achieve conversational understanding of both temporal visual scenes and auditory signals.

Listen

Large Language Models have demonstrated strong capabilities in conversational tasks, but real-world human-computer interaction relies heavily on multimodal communication. While recent systems can process text alongside static images or isolated audio signals, they fail to comprehensively understand video, which requires tracking visual changes over time while simultaneously interpreting auditory information.

The article sets out to design and demonstrate Video-LLaMA, a multimodal framework that enables large language models to understand both visual dynamics and auditory signals in videos for conversational human-computer interaction.

To achieve this efficiently without retraining the core language model, the approach freezes the underlying text model and pre-trained perception encoders, using dedicated visual and audio adapter modules to translate multimodal inputs into language-compatible representations. The visual branch incorporates a temporal position embedding layer and a video transformer adapter to aggregate frame sequences, pre-trained on video datasets like WebVid-2M and fine-tuned on instruction-following datasets. The audio branch leverages a universal multimodal encoder and an audio transformer adapter. Because paired audio-text training data is scarce, the audio branch is trained using visual-text datasets, relying on the encoder's shared multimodal space to achieve zero-shot audio understanding during inference.

Key findings show that Video-LLaMA successfully bridges video perception and text generation. First, the model achieves simultaneous audiovisual comprehension, accurately answering questions about both background sounds and visual elements within the same video. Second, it effectively captures temporal dynamics, correctly identifying sequential human actions and moving objects across frames. Third, the system demonstrates strong visual reasoning and static image understanding, identifying unusual elements in scenes and recognizing famous landmarks and public figures.

These findings prove that multi-branch adapter training is a viable, parameter-efficient pathway to build unified audio-visual conversational assistants without expensive full-model retraining. However, the system currently operates as an early-stage prototype with notable limitations, including perceptual errors driven by limited training dataset scale, high computational overhead when processing long videos like television shows, and the tendency to generate inaccurate or hallucinated text inherited from the base language model.

For future development and practical application, the article highlights the need to construct larger, higher-quality audio-video-text datasets and develop more efficient architectures capable of handling long-form video content before deploying such models into high-stakes, production-grade environments.

Cover for Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Abstract

We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video. Video-LLaMA bootstraps cross-modal training from the frozen pre-trained visual and audio encoders and the frozen LLMs. Unlike previous works that complement LLMs to process the visual or audio signals only, Video-LLaMA enables video comprehension by tackling two challenges: (1) capturing the temporal changes in visual scenes, (2) integrating audio-visual signals. To counter the first challenge, we propose a Video Q-former to assemble a pre-trained image encoder into our video encoder and introduce a video-to-text generation task to learn video-language correspondence. For the second challenge, we leverage ImageBind, a universal embedding model aligning multiple modalities, as the pre-trained audio encoder and introduce an Audio Q-former on top of ImageBind to learn reasonable auditory query embeddings for the LLM module. To align the output of both visual and audio encoders with LLM's embedding space, we first train Video-LLaMA on massive video/image-caption pairs and then tune our model with visual-instruction datasets of moderate amount but higher quality. We found Video-LLaMA shows the ability to perceive and comprehend video content and generate meaningful responses grounded in the visual and auditory information presented in the videos.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Architecture
  • 2.1.1 Vision-Language Branch
  • 2.1.2 Audio-Language Branch
  • 2.2 Multi-branch Cross-Modal Training
  • 2.2.1 Training of Vision-Language Branch
  • 2.2.2 Training of Audio-Language Branch
  • 3 Related Works
  • 4 Examples
  • 5 Conclusion
  • 6 Limitations
  • References
  • A Appendix

Knowls

  1. Knowl 1 — Overall Architecture of Video-LLaMA

    model/method

    Video-LLaMA is a multi-modal framework designed to enable frozen Large Language Models (LLMs, such as LLaMA or Vicuna) to understand and converse about both visual and auditory content in videos.

    The framework adopts a multi-branch architecture consisting of two parallel pathways:

    1. Vision-Language Branch: Transforms input video frames into a sequence of visual query vectors.
    2. Audio-Language Branch: Transforms input audio signals into a sequence of auditory query vectors.

    Both branches project their respective modality features into the textual embedding space of the LLM via dedicated Query Transformers (Q-Formers) and linear projection layers. The resulting visual and auditory query vectors are concatenated with text embeddings as multi-modal soft prompts, providing context that guides the frozen LLM in generating text responses grounded in both video frames and audio.

  2. Knowl 2 — Vision-Language Branch Architecture

    model/method

    The Vision-Language Branch of Video-LLaMA processes visual video inputs through four components:

    1. Frozen Visual Encoder: Uses the visual backbone of BLIP-2, comprising a ViT-G/14 from EVA-CLIP and a pre-trained Q-Former. Given an input video of NN frames, the encoder extracts frame-level representations V=[v1,v2,…,vN]V = [v_1, v_2, \dots, v_N], where each vi∈RKf×dfv_i \in \mathbb{R}^{K_f \times d_f} represents KfK_f feature embeddings of dimension dfd_f for the ii-th frame.
    2. Position Embedding Layer: Applies learnable positional embeddings across the frame representations v1,…,vNv_1, \dots, v_N to inject temporal sequence information that is missing from independently computed frame features.
    3. Video Q-Former: Shares the Query Transformer architecture from BLIP-2. It takes the position-encoded frame representations as input and aggregates them into a fixed-size video representation v^∈RkV×dv\hat{v} \in \mathbb{R}^{k_V \times d_v}, containing kVk_V video embedding vectors of dimension dvd_v.
    4. Linear Projection Layer: Maps the kVk_V video embeddings from dimension dvd_v to the token embedding dimension of the target LLM. These projected vectors serve as visual soft prompts prepended to the user query text embeddings.
  3. Knowl 3 — Audio-Language Branch Architecture

    model/method

    The Audio-Language Branch of Video-LLaMA processes audio signals extracted from video through the following sequence:

    1. Audio Sampling and Preprocessing: Uniformly samples MM segments of 2-second audio clips from the video and converts each clip into a spectrogram using 128 mel-spectrogram bins.
    2. Frozen Audio Encoder: Uses the pre-trained ImageBind audio encoder to map each 2-second spectrogram into a dense feature vector, resulting in an audio representation sequence A=[a1,a2,…,aM]A = [a_1, a_2, \dots, a_M].
    3. Position Embedding Layer: Injects temporal order across audio segments by adding learnable positional embeddings to a1,…,aMa_1, \dots, a_M.
    4. Audio Q-Former: Uses a Query Transformer architecture identical to BLIP-2's Q-Former. It interacts with the position-encoded audio segment sequence to output a fixed-length audio feature sequence A^∈RKa×da\hat{A} \in \mathbb{R}^{K_a \times d_a}, where KaK_a is the number of learned audio queries and dad_a is the query feature dimension.
    5. Linear Projection Layer: Projects the audio embeddings into the LLM's text embedding dimension. The projected auditory vectors are concatenated with text embeddings alongside visual query vectors to guide the LLM.
  4. Knowl 4 — Two-Stage Training of the Vision-Language Branch

    model/method

    The Vision-Language Branch in Video-LLaMA is trained in two sequential stages with the vision encoder and LLM kept frozen:

    1. Stage 1: Large-Scale Vision-Language Pre-training:

      • Data: WebVid-2M (2 million short video clips with stock footage descriptions) and CC595k (a filtered subset of CC3M image-caption pairs, where each image is treated as a single-frame video).
      • Task: Video-to-text generation, where the learnable position embeddings, Video Q-Former, and linear projection layer are trained to prompt the LLM to generate descriptive text given the visual embedding prefix.
      • Objective: Maximize visual knowledge capture across video and static image semantics.
    2. Stage 2: Visual Instruction Fine-Tuning:

      • Data: High-quality visual instruction datasets, including the image detail description dataset from MiniGPT-4, the image-instruction dataset from LLaVA, and the video-instruction dataset from Video-Chat.
      • Objective: Recover instruction-following capability and enhance detailed conversational grounding for both static images and videos.
  5. Knowl 5 — Zero-Shot Audio Alignment via ImageBind Common Embedding Space

    model/method

    Directly training the Audio-Language Branch on audio-text datasets is constrained by the scarcity of large-scale, paired audio-caption data. Video-LLaMA bypasses this bottleneck by exploiting the multi-modal alignment property of ImageBind:

    1. ImageBind maps multiple modalities (including vision and audio) into a single, shared embedding space.
    2. The learnable parameters of the Audio-Language Branch (the positional embeddings, Audio Q-Former, and linear layer) are trained using visual-text datasets rather than audio-text datasets, following the identical training pipeline as the vision branch.
    3. By learning to map ImageBind's shared representation space into the LLM token embedding space via visual-text supervision, the Audio-Language Branch automatically gains the ability to interpret audio inputs mapped into the same space by ImageBind.
    4. Consequently, Video-LLaMA exhibits zero-shot audio understanding during inference without having been trained directly on audio-text pairs.
  6. Knowl 6 — Modality Support Comparison Across Multi-Modal LLMs

    data/table

    Existing multi-modal large language models typically support either static images, silent video, or audio, but not all three simultaneously within an end-to-end architecture.

    Model Name Static Image Silent Video Audio
    BLIP-2 ✓ ×\times ×\times
    MiniGPT-4 ✓ ×\times ×\times
    LLaVA ✓ ×\times ×\times
    mPLUG-Owl ✓ ✓ ×\times
    VideoChat ✓ ✓ ×\times
    AudioGPT ×\times ×\times ✓
    Video-ChatGPT ✓ ✓ ×\times
    Video-LLaMA ✓ ✓ ✓

    Video-LLaMA uniquely integrates simultaneous visual and auditory comprehension alongside static image processing into a single end-to-end framework, contrasting with models like VideoChat and Video-ChatGPT that discard audio, or AudioGPT that lacks video reasoning.

  7. Knowl 7 — Multi-Modal Capabilities and Grounded Conversational Abilities

    empirical result

    Qualitative evaluation of Video-LLaMA across video, audio, and image grounded conversations demonstrates four key capabilities:

    1. Audio-Visual Joint Perception: Video-LLaMA can answer distinct queries targeting visual content (e.g., identifying that a person is playing a saxophone or wearing glasses) and acoustic content (e.g., detecting background applause, footsteps, or a dog barking) from the same video.
    2. Temporal Dynamics Reasoning: The model identifies actions over time and movement directions (e.g., describing the trajectory of a moving boat or sequence of human gestures).
    3. Static Image Understanding: The model can describe complex scenes and identify unusual anomalies (e.g., a person ironing clothes on top of a car).
    4. Common-Sense and Concept Recognition: The model identifies famous cultural landmarks (e.g., the United States Capitol) and media characters (e.g., Jon Snow and Daenerys Targaryen from Game of Thrones) and reasons about their relationships.
  8. Knowl 8 — Limitations of Video-LLaMA

    limitation

    Video-LLaMA exhibits three primary limitations:

    1. Perception Capacities: Fine-grained perception is constrained by the scale and quality of existing open-source audio-video-text alignment datasets.
    2. Long Video Processing: The architecture has limited capacity to process extended videos (such as full TV episodes or movies) due to computational and memory constraints associated with dense frame and audio sampling.
    3. Hallucination: Video-LLaMA inherits the hallucination tendencies of its underlying pre-trained frozen language model, occasionally generating ungrounded facts or assertions.

Coverage note — None was omitted; the knowls cover the complete system architecture, training strategies for both branches, modality alignment methodology, qualitative results, and identified limitations.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022a. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bińkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022b. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  4. 4.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In IEEE International Conference on Computer Vision.
  5. 5.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  8. 8.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2022. Eva: Exploring the limits of masked visual representation learning at scale. arXiv preprint arXiv:2211.07636.
  9. 9.Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, W. Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, and Yu Jiao Qiao. 2023. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010.
  10. 10.Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15180–15190.
  11. 11.Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2023a. Audiogpt: Understanding and generating speech, music, sound, and talking head. arXiv preprint arXiv:2304.12995.
  12. 12.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. 2023b. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045.
  13. 13.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023a. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726.
  14. 14.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  15. 15.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wen Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023c. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355.
  16. 16.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  17. 17.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Ming-Hui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207.
  18. 18.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424.
  19. 19.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  20. 20.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman ´ Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  21. 21.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565. Association for Computational Linguistics.
  22. 22.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugging-gpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  23. 23.Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:2305.16355.
  24. 24.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2023. Generative pretraining in multimodality. arXiv preprint arXiv:2307.05222.
  25. 25.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  26. 26.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212.
  27. 27.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning.
  28. 28.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023a. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671.
  29. 29.Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yilun Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, and Yu Wu. 2023b. On decoder-only architecture for speech-to-text and large language model integration. arXiv preprint arXiv:2307.03917.
  30. 30.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023a. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. arXiv preprint arXiv:2304.01196.
  31. 31.Haiyang Xu, Qinghao Ye, Mingshi Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qiuchen Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Feiran Huang, and Jingren Zhou. 2023b. mplug-2: A modularized multi-modal foundation model across text, image and video. arXiv preprint arXiv:2302.00402.
  32. 32.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yi Zhou, Junyan Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qiang Qi, Ji Chao Zhang, and Feiyan Huang. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178.
  33. 33.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549.
  34. 34.Ao Zhang, Hao Fei, Yuan Yao, Wei Ji, Li Li, Zhiyuan Liu, and Tat-Seng Chua. 2023a. Transfer visual prompt generator across llms. arXiv preprint arXiv:23045.01278.
  35. 35.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Peng Peng Wang, Yaqian Zhou, and Xipeng Qiu. 2023b. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000.
  36. 36.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  37. 37.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Zhang, H., et al. “Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding”. arXiv, 2023, http://arxiv.org/abs/2306.02858v4.
APA
Zhang, H., Li, X., & Bing, L. (2023). Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. arXiv. http://arxiv.org/abs/2306.02858v4
Chicago
Zhang, H., X. Li, and L. Bing. 2023. “Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding”. arXiv. http://arxiv.org/abs/2306.02858v4.
Harvard
Zhang, H., Li, X. and Bing, L. (2023) “Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.02858v4.
Vancouver
1. Zhang H, Li X, Bing L (2023) Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. arXiv

BibTeX

@article{zhang2023video,
  title = {Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding},
  author = {Zhang, Hang and Li, Xin and Bing, Lidong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.02858v4},
  eprint = {2306.02858}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/