VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Haoxin ChenYong ZhangXiaodong CunMenghan XiaXintao WangChao WengYing Shan

article2024CVPR579 citations

Presents a training strategy that enables text-to-video diffusion models to overcome low-quality video datasets by finetuning spatial modules with high-quality images without sacrificing motion dynamics.

Listen

Text-to-video artificial intelligence models have advanced rapidly, yet leading commercial systems rely on massive proprietary datasets of well-curated, high-resolution videos that remain inaccessible to the broader research community. In contrast, publicly available video datasets such as WebVid-10M are restricted by low visual resolution, compression artifacts, and embedded watermarks. Consequently, non-commercial models directly trained on open data inherit these visual flaws, creating a substantial quality gap between proprietary platforms and accessible research tools.

The article demonstrates that high-performance, generic text-to-video models can be trained without any high-quality video footage by decoupling motion learning from visual appearance learning. The researchers propose a training framework that leverages low-quality public videos exclusively to establish temporal motion consistency, followed by direct refinement using high-resolution synthetic images to instill sharpness, aesthetic quality, and complex concept composition.

To achieve this, the authors evaluated how video models process spatial appearance and temporal movement when built upon existing image diffusion backbones. They conducted comparative experiments analyzing full-parameter training—where all spatial and temporal layers are updated simultaneously—against partial-parameter training, where spatial layers remain frozen. The models were trained on 10 million low-quality video clips from WebVid-10M paired with LAION-COCO image data on 32 graphics processing units, and then perturbed or fine-tuned using the JourneyDB dataset comprising 4 million high-resolution synthetic images. Performance was benchmarked through standardized objective metrics alongside blinded human expert evaluations.

The article establishes several key findings. First, fully training both spatial and temporal modules creates a significantly stronger coupling between appearance and motion, allowing the model to withstand substantial subsequent modifications without motion breakdown; in contrast, partially trained models quickly lose temporal coherence and freeze into static frames. Second, directly fine-tuning only the spatial modules with high-quality synthetic images significantly elevates visual quality (achieving an aesthetic score of 82.57 compared to 46.55 in the base model) while eliminating artifacts such as watermarks without degrading motion smoothness. Third, synthetic images generated from advanced text-to-image tools provide superior concept composition compared to standard web-scraped image collections. Finally, across objective benchmarks and user evaluations, the resulting model—VideoCrafter2—matches or surpasses leading open-source models and attains visual quality competitive with proprietary commercial systems, winning a 61% user preference margin over its direct predecessor.

These results demonstrate that organizations do not need to undertake prohibitive copyright, storage, and processing costs to collect massive proprietary video repositories for generative video modeling. Instead, visual fidelity and motion dynamics can be disentangled at the data level. However, the authors note clear operational boundaries: while visual quality and text alignment match top benchmarks, motion quality still trails behind systems trained on significantly larger, proprietary video volumes. Organizations developing generative video pipelines should adopt full-parameter base training followed by spatial-only refinement using high-quality synthetic imagery, while focusing subsequent research on scaling temporal diversity to close the remaining motion gap.

No sufficiently relevant recommendations were found.

Cover for VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Abstract

Text-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with minimal noise, excellent details, and high aesthetic scores. However, these models rely on large-scale, well-filtered, high-quality videos that are not accessible to the community. Many existing research works, which train models using the low-quality WebVid-10M dataset, struggle to generate high-quality videos because the models are optimized to fit WebVid-10M. In this work, we explore the training scheme of video models extended from Stable Diffusion and investigate the feasibility of leveraging low-quality videos and synthesized high-quality images to obtain a high-quality video model. We first analyze the connection between the spatial and temporal modules of video models and the distribution shift to low-quality videos. We observe that full training of all modules results in a stronger coupling between spatial and temporal modules than only training temporal modules. Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model. Evaluations are conducted to demonstrate the superiority of the proposed method, particularly in picture quality, motion, and concept composition.

Citation

MLA
Chen, H., et al. “VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models”. arXiv, 2024, http://arxiv.org/abs/2401.09047v1.
APA
Chen, H., Zhang, Y., Cun, X., Xia, M., Wang, X., Weng, C., & Shan, Y. (2024). VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv. http://arxiv.org/abs/2401.09047v1
Chicago
Chen, H., Y. Zhang, X. Cun, et al. 2024. “VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models”. arXiv. http://arxiv.org/abs/2401.09047v1.
Harvard
Chen, H. et al. (2024) “VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.09047v1.
Vancouver
1. Chen H, Zhang Y, Cun X, Xia M, Wang X, Weng C, Shan Y (2024) VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv

BibTeX

@article{chen2024videocrafter2,
  title = {VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models},
  author = {Chen, Haoxin and Zhang, Yong and Cun, Xiaodong and Xia, Menghan and Wang, Xintao and Weng, Chao and Shan, Ying},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.09047v1},
  eprint = {2401.09047}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE