T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Kaiyue SunKaiyi HuangXian LiuYue WuZihan XuZhenguo LiXihui Liu
Presents the first systematic benchmark and multi-modal evaluation suite for compositional text-to-video generation across seven spatial and temporal categories, revealing critical limitations in current video generation models.
Text-to-video generative models have advanced rapidly, yet existing evaluation standards mainly test simple, single-object prompts and overall video quality. In real-world applications, generative systems must faithfully synthesize complex, dynamic scenes where multiple objects, specific attributes, distinct movements, and precise interactions occur simultaneously over time. The article addresses this critical evaluation gap by systematically assessing how well current models handle compositional text-to-video generation.
The main objective of the article is to establish a standardized benchmark, named T2V-CompBench, and develop targeted evaluation metrics to assess the compositional generation capabilities of text-to-video models across diverse spatio-temporal requirements.
To accomplish this, the authors designed a prompt suite consisting of 1,400 text prompts divided equally across seven distinct categories: consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. These prompts were constructed using high-frequency vocabulary from real-world user data and GPT-4 generation, ensuring every prompt includes temporal motion with active verbs. To evaluate video alignment accurately across time, the authors devised a multi-pronged automated framework employing multimodal large language models (such as Grid-LLaVA and D-LLaVA), object detection models (GroundingDINO with depth estimation), and point-tracking algorithms (DOT). The validity of these automated metrics was verified by evaluating their statistical correlation against 651 human-annotated video ratings across 23 different text-to-video models, including 17 open-source and 6 commercial systems.
The investigation produced several key findings regarding model capabilities. First, dynamic attribute binding—where an object's appearance changes over time—is the single most difficult task; most systems score near zero because they generate static attributes and ignore temporal transitions. Second, models struggle significantly with spatial positioning, motion direction, and generative numeracy. Current systems frequently confuse directional terms like left and right, fail to generate directed movement against moving camera backgrounds, and rarely produce correct counts when asked for more than three objects. Third, while models perform comparatively better in consistent attribute binding, action binding, and object interactions, they still regularly suffer from attribute misassignment, missed secondary objects, or static outputs. Overall, top-performing commercial models like PixVerse-V3 generally outperform open-source models, but compositional fidelity remains low across the board.
These findings indicate that while today's video generation models produce high-resolution, visually appealing single frames, they lack the spatial reasoning and temporal control necessary for production-grade use cases requiring strict prompt adherence. For stakeholders, deploying current models in settings that demand precise visual storytelling, dynamic physical interactions, or exact object counts poses significant operational and quality risks, as automated adherence to complex instructions is unreliable.
To address these deficiencies, development teams and researchers should focus training efforts on temporal dynamics and explicit motion control rather than purely optimizing frame-level visual aesthetics. Specifically, model developers should invest in building specialized video datasets with granular, dynamic captions and explore architecture designs that incorporate explicit layout planning and motion guidance modules. In the interim, evaluators should adopt multi-method evaluation suites—combining language models, object detection, and point tracking—over traditional image-based similarity metrics like CLIP.
The conclusions of this article are supported by strong human-correlation validation across a wide range of state-of-the-art models. However, readers should consider certain boundary conditions: the benchmark primarily evaluates short video clips (2 to 5 seconds) focusing on discrete objects with well-defined physical boundaries rather than continuous, boundaryless visual elements. As longer video generation evolves, further benchmark extensions will be required to assess sustained multi-scene compositionality.
- Paper: T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. (2023). This benchmark establishes the foundational compositional evaluation categories and metric methodologies for text-to-image synthesis that T2V-CompBench directly adapts and extends into dynamic video generation.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). VBench defines the standard multidimensional evaluation framework for video generative models, providing crucial context for how T2V-CompBench identifies and isolates compositional limitations.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench demonstrates how to construct fine-grained, dynamic temporal perception benchmarks using multimodal models, inspiring the temporal evaluation paradigms used in T2V-CompBench.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench introduces robust multimodal evaluation principles and fine-grained capability breakdowns that underpin modern multimodal model-based evaluation pipelines.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). MIGC provides an essential exploration of multi-instance attribute binding and spatial leakage failure modes that motivate the multi-object compositional criteria in T2V-CompBench.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). Imagen Video establishes foundational cascaded diffusion architectures for text-to-video generation, providing key technical background on the base generative systems evaluated in the benchmark.
- Paper: Towards Accurate Generative Models of Video: A New Metric & Challenges, Thomas Unterthiner et al. (2018). This work introduces Fréchet Video Distance (FVD), illustrating the traditional holistic video quality metrics whose limitations necessitate granular compositional benchmarks like T2V-CompBench.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). WorldSimBench builds on compositional video generation evaluation by measuring how faithfully generative video models act as physically compliant world simulators in interactive environments.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). SpatialGenEval extends compositional spatial evaluation to detailed scenes and spatial intelligence reasoning across complex prompts.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). CogVideoX introduces advanced Diffusion Transformer architectures designed to overcome the semantic and action binding challenges highlighted by compositional benchmarks like T2V-CompBench.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). Tora incorporates explicit trajectory-guided motion conditioning into video diffusion transformers, addressing the dynamic motion and action binding deficiencies exposed in T2V-CompBench.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME evaluates the dynamic video reasoning capabilities of multimodal LLMs that serve as core automated judges in video evaluation pipelines.
