VBench: Comprehensive Benchmark Suite for Video Generative Models
Ziqi HuangYinan HeJiashuo YuFan ZhangChenyang SiYuming JiangYuanhan ZhangTianxing WuQin JinNattapol Chanpaisit
Presents VBench, an evaluation suite that measures video generation quality across 16 human-aligned dimensions—including temporal consistency, motion smoothness, and subject identity—to provide precise diagnostic benchmarks for video synthesis models.
Artificial intelligence models for generating video from text descriptions have advanced rapidly, yet evaluating their actual performance remains a critical bottleneck. Traditional automated metrics often compress complex video behavior into a single score that fails to align with human perception, while standard video quality assessment tools do not account for generative artifacts like subject morphing or temporal flickering. To resolve these challenges, the article develops and validates VBench, an open-source, comprehensive evaluation benchmark suite designed to systematically dissect video generation quality into granular, objective, and human-aligned criteria.
The benchmark decomposes video evaluation into 16 distinct dimensions structured under two core categories: Video Quality, which assesses intrinsic visual aspects such as subject consistency, background consistency, motion smoothness, temporal flickering, dynamic motion degree, and aesthetic quality without prompt context; and Video-Condition Consistency, which evaluates how accurately generated content reflects prompt semantics (e.g., object classes, human actions, spatial relationships) and requested visual styles. To evaluate text-to-video systems, the researchers created specialized prompt suites encompassing roughly 100 targeted test cases per dimension and 800 prompts across eight diverse content categories (such as humans, animals, scenery, and food). Automated evaluation pipelines were benchmarked across four major open-source text-to-video models (LaVie, ModelScope, VideoCrafter, and CogVideo) alongside text-to-image baselines. To ensure validity, automated scores were verified against large-scale human preference annotations across all dimensions.
The analysis produced several critical findings regarding current generative models. First, VBench automated evaluations demonstrate strong statistical correlation with human perceptual preferences across every dimension, reaching Spearman correlation coefficients between 80% and nearly 100% across most categories. Second, the evaluation uncovered a significant trade-off between temporal consistency and motion dynamics: models that achieved high scores in subject and background stability (such as LaVie) frequently generated relatively static scenes, whereas models capable of intense dynamic motion (such as VideoCrafter) suffered sharp declines in temporal consistency. Third, text-to-video models severely lag behind modern text-to-image models in compositionality; video generators scored between 18% and 39% on generating multiple objects and between 18% and 37% on spatial relationships, whereas leading image models reached approximately 70% and 86% respectively. Finally, performance varied substantially by subject matter: categories involving complex, articulated motion (such as humans and vehicles) exhibited the lowest quality across all dimensions despite human-focused data making up the largest share (26%) of standard training datasets like WebVid-10M, while curated categories like food consistently yielded high aesthetic scores despite comprising only 11% of the data.
These findings show that simply scaling raw dataset volume will not resolve core generative deficiencies in complex scenes or multi-object composition. The documented trade-offs indicate that single-number evaluation metrics risk incentivizing models that generate nearly static outputs to artificially inflate visual quality scores. For organizations investing in or building video models, the results demonstrate that technical roadmaps must prioritize curated, high-quality training data, upgraded text encoders, and specialized intermediate controls (such as skeletal motion or spatial bounding boxes) rather than relying exclusively on larger compute and data scale.
Based on these insights, developers and stakeholders should adopt multi-dimensional benchmarking to evaluate models based on specific deployment needs (e.g., prioritizing spatial accuracy for e-commerce or motion dynamics for animation). Dataset curators should utilize fine-grained quality dimensions to systematically filter training data. Future research efforts should focus on expanding benchmark coverage to image-to-video and video editing workflows while adding safety, fairness, and content integrity dimensions. Because the current study was limited to four open-source text-to-video models and 2-second clips, decision-makers can rely with high confidence on the benchmark's dimensional accuracy, but should exercise measured caution when generalizing performance observations to longer video durations or emerging closed-source commercial systems.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It introduces foundational latent video diffusion models and standard spatiotemporal generation baselines that VBench specifically targets and benchmarks.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It establishes the core video diffusion framework and standard evaluation metrics like FVD whose limitations motivate VBench's multi-dimensional approach.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). It outlines high-definition cascaded text-to-video generation paradigms and earlier evaluation protocols that contextualize modern video generative assessment.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). It presents autoregressive transformer models for variable-length video generation from text prompts evaluated across temporal consistency dimensions.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It establishes the hierarchical taxonomy and fine-grained multi-skill evaluation principles for multimodal models that inspired comprehensive evaluation suites like VBench.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). It introduces the methodology of massive, multi-task, human-aligned benchmarking suites designed to quantify model capabilities across diverse fine-grained tasks.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It scales latent video diffusion models using curated datasets and rigorous human preference evaluations that build directly on structured video generation benchmarks.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). It extends video generative model evaluation beyond perceptual metrics to embodied simulation and physical actionability.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). It advances open-source text-to-video generation with expert transformers and 3D VAEs, relying on fine-grained video quality dimensions for evaluation.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). It introduces a space-time U-Net architecture designed specifically to address temporal flickering and motion smoothness challenges highlighted by VBench.
- Paper: Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning, Rohit Girdhar et al. (2024). It factorizes video generation via explicit image conditioning to improve motion quality and human-aligned visual consistency.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It provides a comprehensive multimodal evaluation benchmark for video understanding across long contexts, mirroring the multi-dimensional benchmark philosophy of VBench.
