Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

Yang JinZhicheng SunKun XuKun XuLiwei ChenHao JiangQuzhe HuangChengru SongYuliang LiuDi Zhang

article2024ICML94 citations

Introduces an efficient multimodal framework that decomposes videos into keyframes and discrete motion vectors, enabling large language models to jointly comprehend and generate video content with significantly fewer tokens.

Listen

Multimodal artificial intelligence models have achieved remarkable success in processing text and static images, but extending these models to dynamic video remains a significant hurdle. Video contains complex temporal dynamics such as moving objects and camera transitions. Existing artificial intelligence approaches either process frames as isolated images, which ignores motion, or use heavy three-dimensional processing that produces an overwhelming number of data tokens. This creates an unsustainable computational burden that severely restricts the model's ability to handle longer videos efficiently.

The article introduces and evaluates Video-LaVIT, a unified multimodal pre-training framework that allows a single large language model to comprehend and generate text, images, and videos. The primary objective is to demonstrate that decomposing videos into keyframes and temporal motion representations allows efficient, large-scale multimodal learning without compromising visual or temporal understanding.

To achieve this, the approach breaks video shots into static keyframes, which capture core visual semantics, and lightweight motion vectors, which capture temporal movement. The keyframes are converted into discrete tokens using an existing image model, while motion vectors are extracted during standard video decompression and transformed into discrete motion tokens via a dedicated spatiotemporal encoder. During training, video is represented as an alternating sequence of visual and motion tokens, optimized under a unified next-token prediction objective alongside text and images. For generation, a sequential detokenizer uses a conditional diffusion network to reconstruct the keyframe and subsequent video frames. The authors evaluated the system across 13 multimodal benchmarks using public datasets including WebVid-10M, MSVD, MSRVTT, and UCF-101.

The findings show that this decoupled representation achieves state-of-the-art results across both understanding and generation tasks. On zero-shot video question answering, Video-LaVIT outperformed existing baselines, achieving 73.2% accuracy on MSVD-QA and 50.1% on ActivityNet-QA. In long video understanding on the EgoSchema benchmark, the model scored 37.3%, surpassing the 32.1% achieved by prior models that consumed far more frames. In video generation, the model produced high-quality outputs on UCF-101 and MSR-VTT, matching or exceeding systems trained on massive proprietary datasets while requiring far fewer motion tokens. Ablation experiments confirmed that adding explicit motion tokens improved question-answering accuracy by roughly 3 to 6 percentage points compared to frame-only models, while reducing the motion token count from 256 to 135 actually improved benchmark performance.

These results demonstrate that organizations can train high-performing, dual-capability video understanding and generation systems with substantially reduced computational requirements. By inheriting knowledge from pre-trained image models and encoding only incremental motion, developers avoid the high financial and hardware costs of training massive video architectures from scratch. Furthermore, relying entirely on open, publicly available training datasets provides transparency and simplifies copyright and compliance governance compared to black-box proprietary pipelines.

Stakeholders and engineering teams aiming to deploy multimodal artificial intelligence should adopt decoupled visual-motion architectures to scale video applications efficiently. Before scaling to enterprise production, development teams should run pilot evaluations on domain-specific video streams to assess prompt fidelity and motion stability. Further work is recommended to explore more sophisticated, adaptive keyframe selection methods and to extend context windows for ultra-long video sequences.

Confidence in these findings is high given the consistent performance across standard public benchmarks. However, key limitations remain. The model's context window of 4,096 tokens and the short average duration of the WebVid-10M training data constrain its ability to synthesize very long, scene-shifting videos without repetitive keyframe generation. Decision-makers should also remain cautious regarding typical foundation model risks, including hallucination during comprehension and potential misuse in generating synthetic video content.

Cover for Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

Abstract

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its spatiotemporal dynamics. In this paper, we address such limitations in video-language pre-training with an efficient video decomposition that represents each video as keyframes and temporal motions. These are then adapted to an LLM using well-designed tokenizers that discretize visual and temporal information as a few tokens, thus enabling unified generative pre-training of videos, images, and text. At inference, the generated tokens from the LLM are carefully recovered to the original continuous pixel space to create various video content. Our proposed framework is both capable of comprehending and generating image and video content, as demonstrated by its competitive performance across 13 multimodal benchmarks in image and video understanding and generation. Our code and models are available at https://video-lavit.github.io.

Citation

MLA
Jin, Y., et al. “Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization”. arXiv, 2024, http://arxiv.org/abs/2402.03161v3.
APA
Jin, Y., Sun, Z., Xu, K., Xu, K., Chen, L., Jiang, H., Huang, Q., Song, C., Liu, Y., Zhang, D., Song, Y., Gai, K., & Mu, Y. (2024). Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization. arXiv. http://arxiv.org/abs/2402.03161v3
Chicago
Jin, Y., Z. Sun, K. Xu, et al. 2024. “Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization”. arXiv. http://arxiv.org/abs/2402.03161v3.
Harvard
Jin, Y. et al. (2024) “Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.03161v3.
Vancouver
1. Jin Y, Sun Z, Xu K, et al (2024) Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization. arXiv

BibTeX

@article{jin2024video,
  title = {Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization},
  author = {Jin, Yang and Sun, Zhicheng and Xu, Kun and Xu, Kun and Chen, Liwei and Jiang, Hao and Huang, Quzhe and Song, Chengru and Liu, Yuliang and Zhang, Di and Song, Yang and Gai, Kun and Mu, Yadong},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.03161v3},
  eprint = {2402.03161}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/