VBench: Comprehensive Benchmark Suite for Video Generative Models

Ziqi HuangYinan HeJiashuo YuFan ZhangChenyang SiYuming JiangYuanhan ZhangTianxing WuQin JinNattapol Chanpaisit

article2023CVPR1,914 citationsHighlight

Presents VBench, an evaluation suite that measures video generation quality across 16 human-aligned dimensions—including temporal consistency, motion smoothness, and subject identity—to provide precise diagnostic benchmarks for video synthesis models.

Listen

Artificial intelligence models for generating video from text descriptions have advanced rapidly, yet evaluating their actual performance remains a critical bottleneck. Traditional automated metrics often compress complex video behavior into a single score that fails to align with human perception, while standard video quality assessment tools do not account for generative artifacts like subject morphing or temporal flickering. To resolve these challenges, the article develops and validates VBench, an open-source, comprehensive evaluation benchmark suite designed to systematically dissect video generation quality into granular, objective, and human-aligned criteria.

The benchmark decomposes video evaluation into 16 distinct dimensions structured under two core categories: Video Quality, which assesses intrinsic visual aspects such as subject consistency, background consistency, motion smoothness, temporal flickering, dynamic motion degree, and aesthetic quality without prompt context; and Video-Condition Consistency, which evaluates how accurately generated content reflects prompt semantics (e.g., object classes, human actions, spatial relationships) and requested visual styles. To evaluate text-to-video systems, the researchers created specialized prompt suites encompassing roughly 100 targeted test cases per dimension and 800 prompts across eight diverse content categories (such as humans, animals, scenery, and food). Automated evaluation pipelines were benchmarked across four major open-source text-to-video models (LaVie, ModelScope, VideoCrafter, and CogVideo) alongside text-to-image baselines. To ensure validity, automated scores were verified against large-scale human preference annotations across all dimensions.

The analysis produced several critical findings regarding current generative models. First, VBench automated evaluations demonstrate strong statistical correlation with human perceptual preferences across every dimension, reaching Spearman correlation coefficients between 80% and nearly 100% across most categories. Second, the evaluation uncovered a significant trade-off between temporal consistency and motion dynamics: models that achieved high scores in subject and background stability (such as LaVie) frequently generated relatively static scenes, whereas models capable of intense dynamic motion (such as VideoCrafter) suffered sharp declines in temporal consistency. Third, text-to-video models severely lag behind modern text-to-image models in compositionality; video generators scored between 18% and 39% on generating multiple objects and between 18% and 37% on spatial relationships, whereas leading image models reached approximately 70% and 86% respectively. Finally, performance varied substantially by subject matter: categories involving complex, articulated motion (such as humans and vehicles) exhibited the lowest quality across all dimensions despite human-focused data making up the largest share (26%) of standard training datasets like WebVid-10M, while curated categories like food consistently yielded high aesthetic scores despite comprising only 11% of the data.

These findings show that simply scaling raw dataset volume will not resolve core generative deficiencies in complex scenes or multi-object composition. The documented trade-offs indicate that single-number evaluation metrics risk incentivizing models that generate nearly static outputs to artificially inflate visual quality scores. For organizations investing in or building video models, the results demonstrate that technical roadmaps must prioritize curated, high-quality training data, upgraded text encoders, and specialized intermediate controls (such as skeletal motion or spatial bounding boxes) rather than relying exclusively on larger compute and data scale.

Based on these insights, developers and stakeholders should adopt multi-dimensional benchmarking to evaluate models based on specific deployment needs (e.g., prioritizing spatial accuracy for e-commerce or motion dynamics for animation). Dataset curators should utilize fine-grained quality dimensions to systematically filter training data. Future research efforts should focus on expanding benchmark coverage to image-to-video and video editing workflows while adding safety, fairness, and content integrity dimensions. Because the current study was limited to four open-source text-to-video models and 2-second clips, decision-makers can rely with high confidence on the benchmark's dimensional accuracy, but should exercise measured caution when generalizing performance observations to longer video durations or emerging closed-source commercial systems.

Cover for VBench: Comprehensive Benchmark Suite for Video Generative Models

Abstract

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align with human perceptions; 2) An ideal evaluation system should provide insights to inform future developments of video generation. To this end, we present VBench, a comprehensive benchmark suite that dissects "video generation quality" into specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. VBench has three appealing properties: 1) Comprehensive Dimensions: VBench comprises 16 dimensions in video generation (e.g., subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship, etc). The evaluation metrics with fine-grained levels reveal individual models' strengths and weaknesses. 2) Human Alignment: We also provide a dataset of human preference annotations to validate our benchmarks' alignment with human perception, for each evaluation dimension respectively. 3) Valuable Insights: We look into current models' ability across various evaluation dimensions, and various content types. We also investigate the gaps between video and image generation models. We will open-source VBench, including all prompts, evaluation methods, generated videos, and human preference annotations, and also include more video generation models in VBench to drive forward the field of video generation.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 VBench Suite
  • 3.1 Evaluation Dimension Suite
  • 3.1.1 Video Quality
  • 3.1.2 Video-Condition Consistency
  • 3.2 Prompt Suite
  • 3.3 Human Preference Annotation
  • 4 Experiments
  • 4.1 Per-Dimension Evaluation
  • 4.2 Validating Human Alignment of VBench
  • 4.3 Per-Category Evaluation
  • 4.4 Video Generation V.S. Image Generation
  • 5 Insights and Discussions
  • 6 Conclusion
  • References
  • G More Details on Evaluation Dimension and Method Suite
  • G.1 Video Quality
  • G.2 Video-Condition Consistency
  • H More Details on Prompt Suite
  • H.1 Prompt Suite per Evaluation Dimension
  • H.2 Prompt Suite per Category
  • I Human Preference Annotation
  • I.1 Human Annotation Procedures
  • I.2 VLM Tuning
  • J More Implementation Details
  • J.1 Video Generation Models in Evaluation
  • J.2 Reference Baselines
  • J.3 Normalization for Radar Chart Visualization
  • K Potential Negative Societal Impacts
  • L Limitations and Future Work
  • M Additional Experimental Results

Citation

MLA
Huang, Z., et al. “VBench: Comprehensive Benchmark Suite for Video Generative Models”. arXiv, 2023, http://arxiv.org/abs/2311.17982v1.
APA
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., & Liu, Z. (2023). VBench: Comprehensive Benchmark Suite for Video Generative Models. arXiv. http://arxiv.org/abs/2311.17982v1
Chicago
Huang, Z., Y. He, J. Yu, et al. 2023. “VBench: Comprehensive Benchmark Suite for Video Generative Models”. arXiv. http://arxiv.org/abs/2311.17982v1.
Harvard
Huang, Z. et al. (2023) “VBench: Comprehensive Benchmark Suite for Video Generative Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.17982v1.
Vancouver
1. Huang Z, He Y, Yu J, et al (2023) VBench: Comprehensive Benchmark Suite for Video Generative Models. arXiv

BibTeX

@article{huang2023vbench,
  title = {VBench: Comprehensive Benchmark Suite for Video Generative Models},
  author = {Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and Wang, Yaohui and Chen, Xinyuan and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.17982v1},
  eprint = {2311.17982}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE