VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
Haoxin ChenYong ZhangXiaodong CunMenghan XiaXintao WangChao WengYing Shan
Presents a training strategy that enables text-to-video diffusion models to overcome low-quality video datasets by finetuning spatial modules with high-quality images without sacrificing motion dynamics.
Text-to-video artificial intelligence models have advanced rapidly, yet leading commercial systems rely on massive proprietary datasets of well-curated, high-resolution videos that remain inaccessible to the broader research community. In contrast, publicly available video datasets such as WebVid-10M are restricted by low visual resolution, compression artifacts, and embedded watermarks. Consequently, non-commercial models directly trained on open data inherit these visual flaws, creating a substantial quality gap between proprietary platforms and accessible research tools.
The article demonstrates that high-performance, generic text-to-video models can be trained without any high-quality video footage by decoupling motion learning from visual appearance learning. The researchers propose a training framework that leverages low-quality public videos exclusively to establish temporal motion consistency, followed by direct refinement using high-resolution synthetic images to instill sharpness, aesthetic quality, and complex concept composition.
To achieve this, the authors evaluated how video models process spatial appearance and temporal movement when built upon existing image diffusion backbones. They conducted comparative experiments analyzing full-parameter training—where all spatial and temporal layers are updated simultaneously—against partial-parameter training, where spatial layers remain frozen. The models were trained on 10 million low-quality video clips from WebVid-10M paired with LAION-COCO image data on 32 graphics processing units, and then perturbed or fine-tuned using the JourneyDB dataset comprising 4 million high-resolution synthetic images. Performance was benchmarked through standardized objective metrics alongside blinded human expert evaluations.
The article establishes several key findings. First, fully training both spatial and temporal modules creates a significantly stronger coupling between appearance and motion, allowing the model to withstand substantial subsequent modifications without motion breakdown; in contrast, partially trained models quickly lose temporal coherence and freeze into static frames. Second, directly fine-tuning only the spatial modules with high-quality synthetic images significantly elevates visual quality (achieving an aesthetic score of 82.57 compared to 46.55 in the base model) while eliminating artifacts such as watermarks without degrading motion smoothness. Third, synthetic images generated from advanced text-to-image tools provide superior concept composition compared to standard web-scraped image collections. Finally, across objective benchmarks and user evaluations, the resulting model—VideoCrafter2—matches or surpasses leading open-source models and attains visual quality competitive with proprietary commercial systems, winning a 61% user preference margin over its direct predecessor.
These results demonstrate that organizations do not need to undertake prohibitive copyright, storage, and processing costs to collect massive proprietary video repositories for generative video modeling. Instead, visual fidelity and motion dynamics can be disentangled at the data level. However, the authors note clear operational boundaries: while visual quality and text alignment match top benchmarks, motion quality still trails behind systems trained on significantly larger, proprietary video volumes. Organizations developing generative video pipelines should adopt full-parameter base training followed by spatial-only refinement using high-quality synthetic imagery, while focusing subsequent research on scaling temporal diversity to close the remaining motion gap.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). Stable Video Diffusion establishes the latent-video training and staged refinement context that VideoCrafter2 adapts while changing how low- and high-quality data contribute.
- Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). Make-A-Video introduces the key strategy of learning visual content from image models and motion from unpaired videos, clarifying the lineage of VideoCrafter2’s data-level decoupling.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Align Your Latents explains how pretrained image diffusion models gain temporal layers for video, the architectural starting point for VideoCrafter2’s spatial-versus-temporal training experiments.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Video Diffusion Models provides the foundational diffusion approach to temporal video generation that informs the later trade-offs VideoCrafter2 tests between appearance and motion.
No sufficiently relevant recommendations were found.
