T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Kaiyue SunKaiyi HuangXian LiuYue WuZihan XuZhenguo LiXihui Liu

article2025CVPR194 citations

Presents the first systematic benchmark and multi-modal evaluation suite for compositional text-to-video generation across seven spatial and temporal categories, revealing critical limitations in current video generation models.

Listen

Text-to-video generative models have advanced rapidly, yet existing evaluation standards mainly test simple, single-object prompts and overall video quality. In real-world applications, generative systems must faithfully synthesize complex, dynamic scenes where multiple objects, specific attributes, distinct movements, and precise interactions occur simultaneously over time. The article addresses this critical evaluation gap by systematically assessing how well current models handle compositional text-to-video generation.

The main objective of the article is to establish a standardized benchmark, named T2V-CompBench, and develop targeted evaluation metrics to assess the compositional generation capabilities of text-to-video models across diverse spatio-temporal requirements.

To accomplish this, the authors designed a prompt suite consisting of 1,400 text prompts divided equally across seven distinct categories: consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. These prompts were constructed using high-frequency vocabulary from real-world user data and GPT-4 generation, ensuring every prompt includes temporal motion with active verbs. To evaluate video alignment accurately across time, the authors devised a multi-pronged automated framework employing multimodal large language models (such as Grid-LLaVA and D-LLaVA), object detection models (GroundingDINO with depth estimation), and point-tracking algorithms (DOT). The validity of these automated metrics was verified by evaluating their statistical correlation against 651 human-annotated video ratings across 23 different text-to-video models, including 17 open-source and 6 commercial systems.

The investigation produced several key findings regarding model capabilities. First, dynamic attribute binding—where an object's appearance changes over time—is the single most difficult task; most systems score near zero because they generate static attributes and ignore temporal transitions. Second, models struggle significantly with spatial positioning, motion direction, and generative numeracy. Current systems frequently confuse directional terms like left and right, fail to generate directed movement against moving camera backgrounds, and rarely produce correct counts when asked for more than three objects. Third, while models perform comparatively better in consistent attribute binding, action binding, and object interactions, they still regularly suffer from attribute misassignment, missed secondary objects, or static outputs. Overall, top-performing commercial models like PixVerse-V3 generally outperform open-source models, but compositional fidelity remains low across the board.

These findings indicate that while today's video generation models produce high-resolution, visually appealing single frames, they lack the spatial reasoning and temporal control necessary for production-grade use cases requiring strict prompt adherence. For stakeholders, deploying current models in settings that demand precise visual storytelling, dynamic physical interactions, or exact object counts poses significant operational and quality risks, as automated adherence to complex instructions is unreliable.

To address these deficiencies, development teams and researchers should focus training efforts on temporal dynamics and explicit motion control rather than purely optimizing frame-level visual aesthetics. Specifically, model developers should invest in building specialized video datasets with granular, dynamic captions and explore architecture designs that incorporate explicit layout planning and motion guidance modules. In the interim, evaluators should adopt multi-method evaluation suites—combining language models, object detection, and point tracking—over traditional image-based similarity metrics like CLIP.

The conclusions of this article are supported by strong human-correlation validation across a wide range of state-of-the-art models. However, readers should consider certain boundary conditions: the benchmark primarily evaluates short video clips (2 to 5 seconds) focusing on discrete objects with well-defined physical boundaries rather than continuous, boundaryless visual elements. As longer video generation evolves, further benchmark extensions will be required to assess sustained multi-scene compositionality.

Cover for T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Abstract

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this important ability for evaluation. In this work, we conduct the first systematic study on compositional text-to-video generation. We propose T2V-CompBench, the first benchmark tailored for compositional text-to-video generation. T2V-CompBench encompasses diverse aspects of compositionality, including consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. We further carefully design evaluation metrics of multimodal large language model (MLLM)-based, detection-based, and tracking-based metrics, which can better reflect the compositional text-to-video generation quality of seven proposed categories with 1400 text prompts. The effectiveness of the proposed metrics is verified by correlation with human evaluations. We also benchmark various text-to-video generative models and conduct in-depth analysis across different models and various compositional categories. We find that compositional text-to-video generation is highly challenging for current models, and we hope our attempt could shed light on future research in this direction.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Text-to-video Generation.
  • 2.2. Compositional Text-to-image Generation.
  • 2.3. Benchmarks for Text-to-video Generation.
  • 3. Benchmark Construction
  • 3.1. Problem Definition and Categorization
  • 3.2. Prompt Categories
  • 3.3. Prompt Suite Generation
  • 4. Evaluation Metrics
  • 4.1. MLLM-based Evaluation Metrics
  • 4.2. Detection-based Evaluation Metrics
  • 4.3. Tracking-based Evaluation Metrics
  • 5. Experiments
  • 5.1. Evaluated Text-to-video Models
  • 5.2. Evaluation Metrics
  • 5.3. Human Evaluation Correlation Analysis
  • 5.4. Quantitative Evaluation
  • 5.5. Qualitative Evaluation
  • 6. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — T2V-CompBench benchmark scope and compositionality definition

    definition

    T2V-CompBench is a benchmark for compositional text-to-video generation, where a generated video must jointly realize multiple objects, object attributes, quantities, actions, interactions, and spatial or temporal dynamics specified by a text prompt. The benchmark contains 1,400 prompts, divided equally into seven categories with 200 prompts each. The spatial-composition categories are consistent attribute binding, spatial relationships, and generative numeracy; the temporal-composition categories are dynamic attribute binding, motion binding, action binding, and object interactions. Every prompt contains at least one active verb so that the benchmark evaluates dynamic video generation rather than only static image-like content.

  2. Knowl 2 — Seven compositional prompt categories

    definition

    The seven T2V-CompBench categories test distinct compositional abilities:

    • Consistent attribute binding: two objects have distinct attributes that remain associated with the correct objects throughout the video. The attribute types are color, shape, texture, and human-related attributes; 20% of prompts use challenging or uncommon object–attribute combinations.
    • Dynamic attribute binding: an object's attribute changes over time. The change types are color or illumination, shape or size, texture, and combined changes; 80% of prompts describe common real-world changes and 20% describe uncommon or artificial changes.
    • Spatial relationships: two objects maintain a specified relation across the video: left of, right of, above, below, in front of, or behind. Left/right prompt pairs are constructed contrastively by reversing the relation.
    • Motion binding: one or two objects move in specified directions: leftward, rightward, upward, or downward. Prompts include both single-object motion and two-object cases with opposing directions.
    • Action binding: two objects perform different actions, testing whether each action is assigned to its corresponding object. The category contains 80% common prompts and 20% uncommon prompts, including unusual object coexistence and unusual object–activity pairs.
    • Object interactions: multiple objects undergo dynamic physical interactions that change motion or state, or social interactions between living entities.
    • Generative numeracy: the video must contain the quantity of each object specified by the prompt. Prompts include one object with quantities from one to eight and two-object prompts with separate quantities.
  3. Knowl 3 — Prompt-suite construction from real user vocabulary

    experimental setup

    The prompt suite was built from 1.67 million unique text-to-video prompts collected from Pika Discord channels in the VidProM dataset. WordNet was used to analyze noun and verb metaclasses and frequency distributions. The authors selected frequent, entry-level “thing” nouns with clear object boundaries, such as cars and dogs, rather than vague “stuff” categories such as sky. The resulting vocabulary contained 260 object nouns, 200 active verbs, and 80 attributes, emphasizing verbs describing movement, travel, or physical activity while avoiding static verbs such as think, see, or rest. GPT-4 generated 200 prompts for each category from this vocabulary and returned parsed metadata, including objects, actions, attributes, quantities, and relationships. Human reviewers checked the generated prompts and removed inappropriate or invalid examples.

  4. Knowl 4 — Evaluation design and category-to-metric assignment

    model/method

    T2V-CompBench uses three complementary evaluation families because compositional video quality depends on object identity and attributes within frames as well as motion and temporal changes across frames. For videos of typically 2–5 seconds, the evaluation protocol uniformly extracts 6 frames for multimodal-large-language-model evaluation, extracts 16 frames for detection-based evaluation, and samples at 8 frames per second for tracking-based evaluation. The metric assignment is:

    • Grid-LLaVA evaluates consistent attribute binding, action binding, and object interactions.
    • D-LLaVA evaluates dynamic attribute binding.
    • G-Dino evaluates spatial relationships and generative numeracy.
    • DOT evaluates motion binding.

    The benchmark also compares these metrics with CLIP, BLIP-CLIP, BLIP-BLEU, BLIP-VQA, ViCLIP, PLLaVA, frame-based LLaVA, VPEval-S, and M-GDino.

  5. Knowl 5 — MLLM-based metrics for attributes, actions, and interactions

    model/method

    Grid-LLaVA converts a video into an image grid by uniformly sampling six frames and passing the grid to LLaVA. To reduce hallucination, the evaluator first describes the video and then answers disentangled grading questions using chain-of-thought prompting. For consistent attribute binding, GPT-4 parses the prompt into object–attribute phrases; LLaVA grades the alignment of each phrase with the image grid, and the phrase scores are averaged. For action binding, GPT-4 extracts each object and its associated action; LLaVA checks object presence and scores each object–action pair. For object interactions, LLaVA checks whether the objects are present and evaluates the interaction dynamics, its development, and its outcome.

    D-LLaVA evaluates dynamic attribute binding frame by frame with an image LLM. GPT-4 extracts the initial and final object states, such as “bright green leaf” and “brown leaf.” LLaVA scores every frame against both states, with the scoring function favoring alignment of the first frame with the initial state, the final frame with the final state, and intermediate frames with an appropriate transition between them.

  6. Knowl 6 — Detection- and tracking-based metrics

    model/method

    G-Dino uses GroundingDINO to detect the prompt-specified objects in each sampled frame and removes duplicate detections with high intersection-over-union (IoU). For a 2D relation, let (x1,y1)(x_1,y_1) and (x2,y2)(x_2,y_2) be the centers of the first and second object's detected bounding boxes. The first object is classified as left of the second when x1<x2x_1<x_2 and ∣x1−x2∣>∣y1−y2∣|x_1-x_2|>|y_1-y_2|; analogous sign and axis conditions are used for right, above, and below. The per-frame score is 1−IoU1-\mathrm{IoU} when a detected object pair satisfies the requested relation and 00 otherwise, and the video score is the mean over frames. For front/behind relations, Segment Anything produces object masks and Depth Anything predicts a depth map; an object's depth is the mean depth of pixels within its mask, and the per-frame score uses the detected-pair IoU and relative depth before averaging across frames. For numeracy, each detected object class receives score 11 if its detected count exactly equals the quantity in the prompt and 00 otherwise; class scores are averaged within each frame and then across frames.

    DOT addresses motion binding while separating object motion from camera motion. GroundingSAM extracts foreground-object and background masks, and DOT tracks points in both regions over the video. The mean foreground motion vector minus the mean background motion vector estimates the object's motion relative to the scene. The metric scores whether this relative direction agrees with the direction specified in the prompt.

  7. Knowl 7 — Human correlation validates the proposed metrics

    data/table

    Human evaluation used 15 randomly selected prompts from each category, six text-to-video models, and three annotators per video. This produced 90 generated videos per category; 10 ground-truth videos were additionally included for dynamic attribute binding and 11 for object interactions, for 651 videos in total. The mean human score for each video was correlated with automatic scores using Kendall's τ\tau and Spearman's ρ\rho.

    The proposed metric correlations were:

    Could not parse LaTeX table

    Among the tested metrics, Grid-LLaVA was strongest for consistent attribute binding, action binding, and object interactions; D-LLaVA was strongest for dynamic attribute binding; G-Dino was strongest for spatial relationships and numeracy; and DOT was strongest for motion binding. The results support using category-specific metrics rather than a single generic text–video similarity score.

  8. Knowl 8 — Benchmarking 23 text-to-video models

    experimental setup

    The authors evaluated 23 text-to-video models using their official default implementations: 17 open-source models and 6 commercial models. The evaluated families were diffusion U-Net models, including ModelScope, ZeroScope, LVD, AnimateDiff, MagicTime, Show-1, VideoCrafter2, VideoTetris, Vico, and T2V-Turbo-V2; Diffusion Transformer models, including Latte, Open-Sora 1.1 and 1.2, Open-Sora-Plan v1.0.0 and v1.3.0, CogVideoX-5B, and Mochi; and commercial systems Pika-1.0, Gen-2, Gen-3, Dreamina 1.2, PixVerse-V3, and Kling-1.0. All reported scores are normalized to [0,1][0,1], with higher values indicating better compositional performance.

    The complete benchmark scores were:

    Could not parse LaTeX table
  9. Knowl 9 — Model-family and adaptation trends

    empirical result

    The benchmark reveals a shift in text-to-video model strengths. Earlier systems tend to prioritize single-frame visual quality, whereas later systems place more emphasis on inter-frame dynamics and motion. VideoCrafter2 and models adapted from it perform strongly on consistent attribute binding, action binding, and object interactions, while models such as CogVideoX-5B and T2V-Turbo-V2 are relatively stronger on motion binding, consistent with their motion-oriented architectures or training guidance.

    Adaptation generally improves compositional scores. VideoTetris, Vico, and T2V-Turbo-V2 improve over the VideoCrafter2 lineage in most categories. MagicTime improves dynamic attribute binding relative to AnimateDiff. LVD improves over ModelScope in nearly all categories, particularly spatial relationships and motion directions, which the authors associate with its LLM-guided layout planning; its dynamic-attribute and numeracy performance remains constrained by its base model.

  10. Knowl 10 — Observed failure modes and category difficulty

    empirical result

    Dynamic attribute binding is the most difficult category for the evaluated models. Models frequently preserve an initial object or attribute instead of producing the required temporal transition, even when the prompt explicitly describes a change. Spatial relationships, motion binding, and numeracy are the next most difficult areas: models often confuse left and right, fail to produce the requested motion direction or substantial object movement, and generate incorrect object counts, especially for quantities larger than three. They perform comparatively better on quantities below three.

    Consistent attribute binding, action binding, and object interactions are generally easier but remain unreliable. Models can assign the same action to both objects instead of binding different actions to different objects, omit one of the specified objects, attach an attribute to the wrong object, or generate a static scene that does not depict the full interaction process. These results indicate that high visual quality alone does not guarantee compositional correctness, particularly for temporal changes, object-specific actions, and multi-object control.

Coverage note — Appendix-level prompt vocabularies, detailed evaluator prompts, model-specific implementation settings, and the paper's separate social-impact discussion were omitted because they support the benchmark rather than adding load-bearing results beyond the ten knowls above.

References

  1. 1.James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving image generation with better captions. https://cdn.openai.com/papers/dall-e-3.pdf, 2023.
  2. 2.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  3. 3.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In CVPR, 2023.
  4. 4.Capcut. Dreamina. https://dreamina.capcut.com/ai-tool/home, 2024.
  5. 5.Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. In ACM Trans. Graph., 2023.
  6. 6.Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models. arXiv preprint arXiv:2401.09047, 2024.
  7. 7.Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis. In ICLR, 2024.
  8. 8.Minghao Chen, Iro Laina, and Andrea Vedaldi. Training-free layout control with cross-attention guidance. In WACV, 2024.
  9. 9.Jaemin Cho, Abhay Zala, and Mohit Bansal. Visual programming for text-to-image generation and evaluation. In NeurIPS, 2023.
  10. 10.Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis. In ICML, 2024.
  11. 11.Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured diffusion guidance for compositional text-to-image synthesis. In ICLR, 2023.
  12. 12.Hanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan, and Peter Wonka. Llm blueprint: Enabling text-to-image generation with complex and detailed prompts. In ICLR, 2024.
  13. 13.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. In ICLR, 2024.
  14. 14.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221, 2022.
  15. 15.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: a reference-free evaluation metric for image captioning. In EMNLP, 2021.
  16. 16.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017.
  17. 17.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  18. 18.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In ICLR, 2023.
  19. 19.hpcaitech. Open-sora: Democratizing efficient video production for all, 2024.
  20. 20.Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. In NeurIPS, 2023.
  21. 21.Kaiyi Huang, Chengqi Duan, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation. In TPAMI, 2025.
  22. 22.Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. In CVPR, 2024.
  23. 23.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In ICCV, 2023.
  24. 24.Wonkyun Kim, Changin Choi, Wonseok Lee, and Wonjong Rhee. An image grid can be worth a video: Zero-shot video question answering using a vlm. arXiv preprint arXiv:2403.18406, 2024.
  25. 25.Yunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha, and Jun-Yan Zhu. Dense text-to-image generation with attention modulation. In ICCV, 2023.
  26. 26.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023.
  27. 27.Kuaishou. Kling. https://kling.kuaishou.com/, 2024.
  28. 28.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  29. 29.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023.
  30. 30.Jiachen Li, Qian Long, Jian Zheng, Xiaofeng Gao, Robinson Piramuthu, Wenhu Chen, and William Yang Wang. T2v-turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design. arXiv preprint arXiv:2410.05677, 2024.
  31. 31.Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. In ICCV, 2023.
  32. 32.Zhiheng Li, Martin Renqiang Min, Kai Li, and Chenliang Xu. Stylet2i: Toward compositional and high-fidelity text-to-image synthesis. In CVPR, 2022.
  33. 33.Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. Llm-grounded video diffusion models. In ICLR, 2023.
  34. 34.Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131, 2024.
  35. 35.Han Lin, Abhay Zala, Jaemin Cho, and Mohit Bansal. Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning. In COLM, 2024.
  36. 36.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, 2024.
  37. 37.Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. In ECCV, 2022.
  38. 38.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In ECCV, 2024.
  39. 39.Xuantong Liu, Tianyang Hu, Wenjia Wang, Kenji Kawaguchi, and Yuan Yao. Referee can play: An alternative approach to conditional generation via model inversion. In ICML, 2024.
  40. 40.Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. In CVPR, 2024.
  41. 41.Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation. In NeurIPS, 2024.
  42. 42.Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. Tempcompass: Do video llms really understand videos? In ACL Findings, 2024.
  43. 43.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In CVPR, 2023.
  44. 44.Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024.
  45. 45.Tuna Han Salih Meral, Enis Simsar, Federico Tombari, and Pinar Yanardag. Conform: Contrast is all you need for high-fidelity text-to-image diffusion models. In CVPR, 2024.
  46. 46.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995.
  47. 47.Guillaume Le Moing, Jean Ponce, and Cordelia Schmid. Dense optical tracking: Connecting the dots. In CVPR, 2024.
  48. 48.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2024.
  49. 49.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  50. 50.Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. Benchmark for compositional text-to-image synthesis. In NeurIPS, 2021.
  51. 51.Maitreya Patel, Changhoon Kim, Sheng Cheng, Chitta Baral, and Yezhou Yang. Eclipse: A resource-efficient text-to-image prior for image generations. In CVPR, 2024.
  52. 52.Pika. Pika. https://www.pika.art, 2024.
  53. 53.PixVerse. Pixverse. https://app.pixverse.ai, 2024.
  54. 54.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  55. 55.Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, and Gal Chechik. Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment. In NeurIPS, 2024.
  56. 56.Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159, 2024.
  57. 57.Runway. Gen-2: Generate novel videos with text, images or video clips. https://research.runwayml.com/gen2, 2024.
  58. 58.Runway. Introducing gen-3 alpha: A new frontier for video generation. https://runwayml.com/blog/introducing-gen-3-alpha/, 2024.
  59. 59.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In NeurIPS, 2016.
  60. 60.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  61. 61.Spencer Sterling. Zeroscope. https://huggingface.co/cerspense/zeroscope_v2_576w, 2023.
  62. 62.Genmo Team. Mochi 1. https://github.com/genmoai/models, 2024.
  63. 63.Ye Tian, Ling Yang, Haotian Yang, Yuan Gao, Yufan Deng, Jingmin Chen, Xintao Wang, Zhaochen Yu, Xin Tao, Pengfei Wan, Di Zhang, and Bin Cui. Videotetris: Towards compositional text-to-video generation. In NeurIPS, 2024.
  64. 64.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018.
  65. 65.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, 2022.
  66. 66.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023.
  67. 67.Ruichen Wang, Zekang Chen, Chen Chen, Jian Ma, Haonan Lu, and Xiaodong Lin. Compositional text-to-image synthesis with attention map control of diffusion models. arXiv preprint arXiv:2305.13921, 2023.
  68. 68.Wenhao Wang and Yi Yang. Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models. In NeurIPS, 2024.
  69. 69.Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. In ICLR, 2023.
  70. 70.Zhenyu Wang, Enze Xie, Aoxue Li, Zhongdao Wang, Xihui Liu, and Zhenguo Li. Divide and conquer: Language models can plan and self-correct for compositional text-to-image generation. arXiv preprint arXiv:2401.15688, 2024.
  71. 71.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, 2022.
  72. 72.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021.
  73. 73.Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nuwa: Visual synthesis pre-training for neural visual world creation. In European conference on computer vision, pages 720–736. Springer, 2022.
  74. 74.Qiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui, Zhe Lin, Yang Zhang, and Shiyu Chang. Harnessing the spatio-temporal attention of diffusion models for high-fidelity text-to-image synthesis. In ICCV, 2023.
  75. 75.Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora. Conceptmix: A compositional image generation benchmark with controllable difficulty. In NeurIPS, 2024.
  76. 76.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024.
  77. 77.Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024.
  78. 78.Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu, Stefano Ermon, and Bin Cui. Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms. In ICML, 2024.
  79. 79.Xingyi Yang and Xinchao Wang. Compositional video generation as flow equalization. arXiv preprint arXiv:2407.06182, 2024.
  80. 80.Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024.
  81. 81.Shenghai Yuan, Jinfa Huang, Yujun Shi, Yongqi Xu, Ruijie Zhu, Bin Lin, Xinhua Cheng, Li Yuan, and Jiebo Luo. Magictime: Time-lapse video generation models as metamorphic simulators. arXiv preprint arXiv:2404.05014, 2024.
  82. 82.Shenghai Yuan, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Ruijie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation. 2024.
  83. 83.David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. IJCV, pages 1–15, 2024.
  84. 84.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022.
  85. 85.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Sun, K., et al. “T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 8406–16, https://doi.org/10.1109/CVPR52734.2025.00787.
APA
Sun, K., Huang, K., Liu, X., Wu, Y., Xu, Z., Li, Z., & Liu, X. (2025). T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8406–8416. https://doi.org/10.1109/CVPR52734.2025.00787
Chicago
Sun, K., K. Huang, X. Liu, et al. 2025. “T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8406–16. https://doi.org/10.1109/CVPR52734.2025.00787.
Harvard
Sun, K. et al. (2025) “T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 8406–8416. Available at: https://doi.org/10.1109/CVPR52734.2025.00787.
Vancouver
1. Sun K, Huang K, Liu X, Wu Y, Xu Z, Li Z, Liu X (2025) T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 8406–8416

BibTeX

@inproceedings{Sun_2025, title={T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00787}, DOI={10.1109/cvpr52734.2025.00787}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Sun, Kaiyue and Huang, Kaiyi and Liu, Xian and Wu, Yue and Xu, Zihan and Li, Zhenguo and Liu, Xihui}, year={2025}, month=June, pages={8406–8416} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE