VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

Wenhao WangYi Yang

article2024NeurIPS148 citations

Introduces VidProM, the first large-scale prompt-gallery dataset containing 1.67 million real-user text-to-video prompts paired with 6.69 million synthesized videos from four diffusion models to advance research in video prompt engineering, efficient generation, and synthetic video detection.

Listen

Recent advances in artificial intelligence have enabled diffusion models to generate realistic video content from written instructions. However, these text-to-video systems depend heavily on the user prompts that guide them, and researchers previously lacked large-scale, real-world prompt datasets to study how users interact with these tools or to benchmark video generation performance effectively.

The article addresses this gap by introducing VidProM, the first large-scale prompt-gallery dataset curated specifically for text-to-video diffusion models. The primary objective is to demonstrate how video prompts fundamentally differ from image prompts and to provide a comprehensive public resource for evaluating models, improving generation efficiency, and developing safety and copyright safeguards.

To construct the dataset, the authors collected 1.67 million unique prompts submitted by real users on public Discord channels between July 2023 and February 2024. Using these prompts, they generated 6.69 million videos across four text-to-video diffusion systems (Pika, VideoCrafter2, Text2Video-Zero, and ModelScope), consuming over 50,000 GPU hours. Each prompt was processed with advanced text embeddings supporting up to 8,192 tokens and annotated with safety scores across six categories of potentially harmful or explicit content. The authors also filtered the collection to isolate approximately 1.04 million semantically unique prompts.

The analysis yielded several critical findings. First, video prompts differ sharply from image prompts: they are significantly longer (with roughly 60,000 prompts exceeding 70 words compared to only 15,000 in image benchmarks), incorporate temporal and dynamic descriptions, and focus heavily on human actions rather than static artistic styles. Second, existing fake-image detectors perform poorly when applied to generated video frames, with diffusion-specific detectors achieving around 49% accuracy (essentially random chance). Third, while direct replication of copyrighted training material by generative models occurs in only a small fraction of outputs (for instance, around 2% in open-source systems and even less in commercial tools), current copy-detection systems fail to reliably flag diffusion-based replications. Finally, the authors demonstrated that fine-tuning language models on VidProM enables effective automated prompt completion for video generation.

These findings indicate that generative video requires specialized engineering, evaluation frameworks, and safety solutions rather than direct adaptations from image tools. Organizations deploying or governing generative video face measurable risks around copyright infringement and deepfake detection, as current visual inspection tools do not generalize to video artifacts. Furthermore, because training data often relies on passive captions rather than user-style prompts, bridging this domain gap is necessary to improve commercial generation quality.

The authors recommend using VidProM to benchmark model performance against realistic user queries and explore prompt-retrieval techniques to reduce computing costs by adapting existing outputs rather than generating videos from scratch. They also suggest developing specialized, end-to-end fake video detectors and pairing automated copy filters with human verification to protect intellectual property. Users should note the dataset's primary limitations: the generated videos are relatively short (between 1.6 and 3.0 seconds) and derived from earlier open-source models rather than the latest frontier systems, meaning future updates will be required to reflect rapidly advancing video quality.

arXiv: 2403.06098
Cover for VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

Abstract

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, along with other text-to-video diffusion models, is highly reliant on prompts, and there is no publicly available dataset that features a study of text-to-video prompts. In this paper, we introduce VidProM, the first large-scale dataset comprising 1.67 Million unique text-to-Video Prompts from real users. Additionally, this dataset includes 6.69 million videos generated by four state-of-the-art diffusion models, alongside some related data. We initially discuss the curation of this large-scale dataset, a process that is both time-consuming and costly. Subsequently, we underscore the need for a new prompt dataset specifically designed for text-to-video generation by illustrating how VidProM differs from DiffusionDB, a large-scale prompt-gallery dataset for image generation. Our extensive and diverse dataset also opens up many exciting new research areas. For instance, we suggest

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Text-to-Video Diffusion Models
  • 2.2 Existing Datasets
  • 3 Curating VidProM
  • 4 The Necessity of Introducing VidProM
  • 4.1 Basic Information
  • 4.2 Semantics of prompts
  • 5 Inspiring New Research
  • 6 Automatic Text-to-video Prompt Completion
  • 7 Conclusion
  • References
  • Checklist
  • A Semantic De-duplication Algorithm
  • B Domain Gap between Video Captions and Text-to-Video Prompts
  • C Benchmarking Existing Fake Image Detection on VidProM
  • D Video Copy Detection for Diffusion Models
  • D.1 Experimental Setup
  • D.2 Observations
  • D.3 Limitations and Future Directions
  • E Data Sheet for VidProM

Knowls

  1. Knowl 1 — VidProM Dataset Specification and Curation

    model/method

    VidProM is a large-scale prompt-gallery dataset designed for text-to-video diffusion research. It comprises 1,672,2431,672,243 unique text-to-video prompts collected from real user submissions in official Pika Discord channels between July 2023 and February 2024.

    Each instance in VidProM contains:

    • A text-to-video prompt extracted from user chat messages (filtering out image-to-video prompts and prompts without corresponding videos).
    • A Universally Unique Identifier (UUID) and a timestamp.
    • A 3072-dimensional prompt embedding generated by OpenAI's text-embedding-3-large model (which supports prompt lengths up to 8192 tokens).
    • Six not-safe-for-work (NSFW) probabilities calculated using the Detoxify model across toxicity, obscenity, identity attack, insult, threat, and sexual explicitness. Prompts with NSFW probabilities greater than 0.2 constitute less than 0.5% of the total dataset.
    • Four generated videos per prompt (6,688,9726,688,972 videos in total, spanning 14,381,289.814,381,289.8 seconds of video generated using approximately 50,63150,631 Nvidia V100 GPU hours across 10 servers with 8 GPUs each). The videos are generated by four diffusion models:
      1. Pika (3.03.0 seconds duration per video),
      2. VideoCrafter2 (1.61.6 seconds duration per video),
      3. Text2Video-Zero (2.02.0 seconds duration per video),
      4. ModelScope (2.02.0 seconds duration per video).
  2. Knowl 2 — Semantic Prompt De-duplication Algorithm (VidProS)

    algorithm

    The semantic de-duplication algorithm constructs the dataset subset VidProS, ensuring that all prompts in the subset are pairwise semantically distinct such that no two prompts have an embedding cosine similarity equal to or exceeding a similarity threshold θ=0.8\theta = 0.8.

    Input: Unique prompts Punique={p1,p2,…,pN}P_{unique} = \{p_1, p_2, \dots, p_N\}, prompt embeddings E∈RN×dE \in \mathbb{R}^{N \times d}, similarity threshold θ\theta
    Output: Semantically unique prompts Punique_semP_{unique\_sem}
    Compute similarity matrix S=E×ETS = E \times E^T and initialize Psimilar={}P_{similar} = \{\}
    for i=1i = 1 to NN do
      for j=i+1j = i + 1 to NN do
        if Si,j>θS_{i,j} > \theta then
          Add {pi,pj}\{p_i, p_j\} to PsimilarP_{similar}
        end if
      end for
    end for
    Initialize Pall={}P_{all} = \{\}
    for each pair {pi,pj}\{p_i, p_j\} in PsimilarP_{similar} do
      if pi∉Pallp_i \notin P_{all} then
        Add pip_i to PallP_{all}
      end if
      if pj∉Pallp_j \notin P_{all} then
        Add pjp_j to PallP_{all}
      end if
    end for
    Initialize Pdel={}P_{del} = \{\}
    for each pair {pi,pj}\{p_i, p_j\} in PsimilarP_{similar} do
      if pi∈Pallp_i \in P_{all} and pj∈Pallp_j \in P_{all} then
        Add pip_i to PdelP_{del}
      end if
    end for
    Initialize Punique_sem={}P_{unique\_sem} = \{\}
    for each pip_i in PuniqueP_{unique} do
      if pi∉Pdelp_i \notin P_{del} then
        Add pip_i to Punique_semP_{unique\_sem}
      end if
    end for
    return Punique_semP_{unique\_sem}

    For N=1,672,243N = 1,672,243 prompts and embedding dimension d=3072d = 3072, computing the pairwise similarity matrix and filtering candidate pairs is executed in approximately 0.6040.604 hours distributed across 8 Nvidia A100 GPUs and 128 CPU cores using Faiss. Applying this algorithm reduces the 1,672,2431,672,243 unique prompts in VidProM to 1,038,8051,038,805 semantically unique prompts in VidProS.

  3. Knowl 3 — Comparative Analysis of VidProM and DiffusionDB

    data/table

    VidProM provides substantial structural and semantic improvements compared to the text-to-image dataset DiffusionDB:

    Aspects Details DiffusionDB VidProM
    Prompts No. of unique prompts 1,819,808 1,672,243
    No. of semantically unique prompts 739,010 1,038,805
    Embedding of prompts OpenAI-CLIP OpenAI-text-embedding-3-large
    Maximum length of prompts 77 tokens 8192 tokens
    Time span Aug 2022 Jul 2023 – Feb 2024
    Images / No. of images/videos ∼\sim14 million images ∼\sim6.69 million videos
    Videos No. of sources 1 4
    Average repetition rate per source ∼\sim8.2 1
    Collection method Web scraping Web scraping + Local generation
    GPU consumption - ∼\sim50,631 V100 GPU hours
    Total seconds - ∼\sim14,381,289.8 seconds

    Although both datasets have a comparable number of raw unique prompts, VidProM contains 40.6% more semantically unique prompts (1,038,8051,038,805 vs. 739,010739,010, re-evaluated using text-embedding-3-large), reflecting greater semantic breadth. Prompts in VidProM are substantially longer and more descriptive (nearly 60,000 prompts exceed 70 words and over 25,000 exceed 100 words, compared to almost zero over 100 words in DiffusionDB). VidProM spans 8 months of real-world user activity (compared to 1 month for DiffusionDB) and generates 4 distinct video outputs per prompt from 4 different diffusion models, resulting in an average prompt repetition rate of 1 per model source compared to 8.2 in DiffusionDB.

  4. Knowl 4 — Semantic Distinctions of Text-to-Video Prompts versus Text-to-Image Prompts and Video Captions

    empirical result

    Text-to-video prompts exhibit fundamental semantic and distributional differences compared to text-to-image prompts (from DiffusionDB) and video captions (from Panda-70M):

    1. Temporal and Dynamic Semantics: Video prompts incorporate three dimensions absent in static image prompts:
      • Time Dimension: Explicit descriptions of scene transitions and temporal action sequences.
      • Dynamic Descriptions: Specification of dynamic actions and verbs (e.g., 'flying', 'working', 'writing') rather than static visual properties.
      • Duration Constraints: Phrases specifying time horizons (e.g., '1-minute', 'a long time').
    2. Prompt Length and Complexity: Video prompts are considerably more complex and verbose to capture spatiotemporal dynamics, mirroring real-world prompts used in large models like Sora (which range from 64 to 95 words).
    3. Topic Distribution Discrepancy: WizMap and t-SNE feature visualizations reveal that user prompt distributions are nearly linearly separable from text-to-image prompts and standard video captions. Text-to-image users frequently generate art and painting styles, whereas text-to-video users focus on everyday human activities (e.g., walking, sports) and popular fiction/superheroes (e.g., Spider-Man, cityscapes).
    4. Domain Gap with Video Captions: Video caption datasets (such as Panda-70M) are dominated by landscape scenes and passive captions, lacking real user generative directives (e.g., 'Generate a video of', '4K', camera motion controls) and human-centric action themes.
  5. Knowl 5 — Benchmark of Fake Image Detectors on Diffusion-Generated Videos

    data/table

    State-of-the-art fake image detectors fail to reliably identify middle frames extracted from videos generated by text-to-video diffusion models when evaluated against 10,000 real video frames from DVSC2023 and 10,000 generated video frames from each of Pika, VideoCrafter2, Text2Video-Zero, and ModelScope.

    Method Accuracy (%) mAP (%)
    Pika VC2 T2VZ MS Avg Pika VC2 T2VZ MS Avg
    CNNSpot 51.17 50.18 49.97 50.31 50.41 54.63 41.12 44.56 46.95 46.82
    FreDect 50.07 54.03 69.88 69.94 60.98 47.82 56.67 75.31 64.15 60.99
    Fusing 50.60 50.07 49.81 51.28 50.44 57.64 41.64 40.51 56.09 48.97
    Gram-Net 84.19 67.42 52.48 50.46 63.64 94.32 80.72 57.73 43.54 69.08
    LGrad 53.73 51.75 41.05 60.22 51.69 54.49 53.21 36.69 66.53 52.73
    LNP 43.48 45.10 47.50 45.21 45.32 44.28 44.08 46.81 39.62 43.70
    DIRE 50.53 49.95 48.96 48.32 49.44 49.21 50.44 44.52 48.64 48.20
    UnivFD 49.41 48.65 49.58 57.43 51.27 48.63 42.36 48.46 70.75 52.55

    Key observations:

    1. Detectors designed for GAN-generated images (e.g., CNNSpot, LGrad) or universal/diffusion-specific image artifacts (e.g., DIRE, UnivFD) perform close to random chance (average accuracy between 45.32% and 51.69%), indicating poor cross-modal and cross-architecture generalization.
    2. Detectors utilizing traditional image processing features, namely global texture statistics (Gram-Net, average accuracy 63.64% and mAP 69.08%) and frequency analysis (FreDect, average accuracy 60.98% and mAP 60.99%), achieve the strongest performance among tested methods.
  6. Knowl 6 — AutoT2VPrompt: Automated Text-to-Video Prompt Completion Model

    model/method

    AutoT2VPrompt is a language model fine-tuned specifically to perform text-to-video prompt completion from short user-supplied prefix texts.

    The model is built by fine-tuning Mistral-7B-v0.1 on the prompt corpus of VidProM. The training is formulated as an autoregressive causal language modeling task using the standard causal language modeling script from Hugging Face Transformers and accelerated with DeepSpeed. The fine-tuning procedure is executed on 8 Nvidia A100 GPUs within 2 hours. The resulting model expands short input prefixes (e.g., 'A cat sitting', 'An underwater world', 'Spiderman') into rich, multi-sentence video prompts that describe dynamic action, visual style, scene composition, and camera movements.

  7. Knowl 7 — Empirical Replication Behavior and Copyright Exposure in Video Diffusion Models

    empirical result

    An empirical evaluation of content replication by text-to-video diffusion models against 1 million reference training videos (from VideoCrafter2 and ModelScope datasets) using the FCPL video copy detection model shows that diffusion models can replicate copyrighted frames and scenes from their training data or existing sources.

    Key findings:

    1. Replicated content represents a small fraction of overall generation; only approximately 2% of videos produced by Text2Video-Zero obtain a maximum normalized copy score greater than 0 (where a score >0> 0 indicates a high probability of replicated content).
    2. The proprietary commercial model Pika exhibits the lowest content replication rate among tested models, whereas the zero-shot open-source model Text2Video-Zero exhibits the highest replication rate.
    3. Commercial video generation can replicate copyrighted artistic works (e.g., generating frames that closely replicate Salvador Dalí's copyrighted painting 'The Persistence of Memory'), demonstrating direct copyright infringement risk in unconstrained diffusion generation.
  8. Knowl 8 — Generalization Failure of Conventional Video Copy Detectors on Diffusion Replications

    empirical result

    Existing video copy detection systems fail to detect many instances of visual replication produced by text-to-video diffusion models. Conventional video copy detection architectures (such as FCPL) are trained on hand-crafted transformations of source videos, including color jitter, grayscale conversion, horizontal flipping, blurring, channel changing, overlays, synthetic stripes, aspect ratio resize cropping, pixel shuffling, and spatial stacking.

    Because diffusion models generate synthetic approximations and generative variations of training content rather than standard photometric or geometric perturbations, copy detection models frequently assign a normalized similarity score ≤0\le 0 to true content replications, causing false negative detections.

  9. Knowl 9 — Limitations of the Generated Video Content in VidProM

    limitation

    The initial video collection in VidProM has several limitations regarding generation quality and temporal length:

    1. Video Duration: The generated videos in the core dataset are short: 3.0 seconds for Pika, 1.6 seconds for VideoCrafter2, 2.0 seconds for Text2Video-Zero, and 2.0 seconds for ModelScope.
    2. Visual Quality and Synthesis Artifacts: The videos are generated using earlier open-source diffusion models and initial commercial systems, resulting in lower visual fidelity, lower resolution, and synthesis artifacts compared to newer frontier video models such as Sora.
    3. Dataset Extensions: To mitigate quality and duration constraints in later revisions, additional subsets of 10,000 720p videos each were generated for StreamingT2V (8-second videos), Open-Sora 1.2 (8-second videos), and CogVideoX-2B (6-second videos).

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.OpenAI. Video generation models as world simulators. https://openai.com/research/video-generation-models-as-world-simulators. Accessed: 2024-03-06.
  2. 2.Pika art. https://pika.art/. Accessed: 2024-03-06.
  3. 3.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023.
  4. 4.Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.
  5. 5.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023.
  6. 6.Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. DiffusionDB: A large-scale prompt gallery dataset for text-to-image generative models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 893–911. Association for Computational Linguistics, July 2023.
  7. 7.Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter1: Open diffusion models for high-quality video generation, 2023.
  8. 8.Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu video: Factorizing text-to-video generation by explicit image conditioning. arXiv preprint arXiv:2311.10709, 2023.
  9. 9.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  10. 10.Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jose Lezama, Jonathan Huang, Rachel Hornung, Hartwig Adam, Hassan Akbari, Yair Alon, Vighnesh Birodkar, et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023.
  11. 11.Morph studio. https://app.morphstudio.com. Accessed: 2024-03-06.
  12. 12.Genie 2024. https://sites.google.com/view/genie-2024/home. Accessed: 2024-03-06.
  13. 13.Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv preprint arXiv:2305.10874, 2023.
  14. 14.Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5036–5045, 2022.
  15. 15.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016.
  16. 16.Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In IEEE International Conference on Computer Vision, 2021.
  17. 17.Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing high-resolution video-language representation with large-scale video transcriptions. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  18. 18.Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. 2024.
  19. 19.Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy, Weidong Cai, and Wayne Wu. Celebv-text: A large-scale facial text-video dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14805–14814, 2023.
  20. 20.Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023.
  21. 21.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. arXiv preprint arXiv:2104.04670, 2021.
  22. 22.Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-jian Jiang, and Alexander Rush. PromptSource: An integrated development environment and repository for natural language prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 93–104. Association for Computational Linguistics, 2022.
  23. 23.Tyrrrz. Discordchatexporter. https://github.com/Tyrrrz/DiscordChatExporter. Accessed: 2024-03-06.
  24. 24.Laura Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020.
  25. 25.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008.
  26. 26.Zijie J. Wang, Fred Hohman, and Duen Horng Chau. WizMap: Scalable Interactive Visualization for Exploring Large Machine Learning Embeddings. arXiv 2306.09328, 2023.
  27. 27.Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Zhao, Weisi Lin, Wynne Hsu, Ying Shan, and Mike Zheng Shou. Towards a better metric for text-to-video generation. arXiv preprint arXiv:2401.07781, 2024.
  28. 28.Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
  29. 29.Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation. Advances in Neural Information Processing Systems, 36, 2023.
  30. 30.Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. 2023.
  31. 31.PKU-Yuan Lab and Tuzhan AI etc. Open-sora-plan, April 2024.
  32. 32.Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. The curse of recursion: Training on generated data makes models forget. arXiv preprint arXiv:2305.17493, 2023.
  33. 33.Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al. Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data. arXiv preprint arXiv:2404.01413, 2024.
  34. 34.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland, 2022. Association for Computational Linguistics.
  35. 35.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671. Association for Computational Linguistics, 2022.
  36. 36.Sam Witteveen and Martin Andrews. Investigating prompt engineering in diffusion models. arXiv preprint arXiv:2211.15462, 2022.
  37. 37.Chang Yu, Junran Peng, Xiangyu Zhu, Zhaoxiang Zhang, Qi Tian, and Zhen Lei. Seek for incantations: Towards accurate text-to-image diffusion synthesis through prompt engineering. arXiv preprint arXiv:2401.06345, 2024.
  38. 38.Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36, 2024.
  39. 39.Yanpeng Sun, Qiang Chen, Jian Wang, Jingdong Wang, and Zechao Li. Exploring effective factors for improving visual in-context learning. arXiv preprint arXiv:2304.04748, 2023.
  40. 40.You-Ming Chang, Chen Yeh, Wei-Chen Chiu, and Ning Yu. Antifakeprompt: Prompt-tuned vision-language models are fake image detectors. arXiv preprint arXiv:2310.17419, 2023.
  41. 41.Chandler Timm Doloriel and Ngai-Man Cheung. Frequency masking for universal deepfake detection. arXiv preprint arXiv:2401.06506, 2024.
  42. 42.Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are surprisingly easy to spot...for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  43. 43.Yan Ju, Shan Jia, Lipeng Ke, Hongfei Xue, Koki Nagano, and Siwei Lyu. Fusing global and local features for generalized ai-synthesized image detection. In 2022 IEEE International Conference on Image Processing (ICIP), pages 3465–3469. IEEE, 2022.
  44. 44.Joel Frank, Thorsten Eisenhofer, Lea Schonherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. Leveraging frequency analysis for deep fake image recognition. In International conference on machine learning, pages 3247–3258. PMLR, 2020.
  45. 45.Bo Liu, Fan Yang, Xiuli Bi, Bin Xiao, Weisheng Li, and Xinbo Gao. Detecting generated images by real images. In European Conference on Computer Vision, pages 95–110. Springer, 2022.
  46. 46.Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei. Learning on gradients: Generalized artifacts representation for gan-generated images detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12105–12114, 2023.
  47. 47.Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. Dire for diffusion-generated image detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22445–22455, October 2023.
  48. 48.Rethinking the up-sampling operations in cnn-based generative network for generalizable deepfake detection. 2024.
  49. 49.Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2023.
  50. 50.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  51. 51.Zhenhua Liu, Feipeng Ma, Tianyi Wang, and Fengyun Rao. A similarity alignment model for video copy segment matching. arXiv preprint arXiv:2305.15679, 2023.
  52. 52.Tianyi Wang, Feipeng Ma, Zhenhua Liu, and Fengyun Rao. A dual-level detection method for video copy detection. arXiv preprint arXiv:2305.12361, 2023.
  53. 53.Wenhao Wang, Yifan Sun, and Yi Yang. Feature-compatible progressive learning for video copy detection. arXiv preprint arXiv:2304.10305, 2023.
  54. 54.Shuhei Yokoo, Peifei Zhu, Junki Ishikawa, and Rintaro Hasegawa. 3rd place solution to meta ai video similarity challenge. arXiv preprint arXiv:2304.11964, 2023.
  55. 55.Succinctly. Text2image prompt generator. https://huggingface.co/succinctly/text2image-prompt-generator, 2023. Hugging Face Model Hub.
  56. 56.MistralAI. Mistral-7b-v0.1. https://huggingface.co/mistralai/Mistral-7B-v0.1, 2023. Hugging Face Model Hub.
  57. 57.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506, 2020.
  58. 58.Roberto Henschel, Levon Khachatryan, Daniil Hayrapetyan, Hayk Poghosyan, Vahram Tadevosyan, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Streamingt2v: Consistent, dynamic, and extendable long video generation from text. arXiv preprint arXiv:2403.14773, 2024.
  59. 59.Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, March 2024.
  60. 60.Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024.
  61. 61.Jeff Johnson, Matthijs Douze, and Herve Jegou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547, 2019.
  62. 62.Abhinav Ramesh Kashyap, Devamanyu Hazarika, Min-Yen Kan, Roger Zimmermann, and Soujanya Poria. So different yet so alike! constrained unsupervised text style transfer. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 416–431. Association for Computational Linguistics, May 2022.
  63. 63.Leo Laugier, John Pavlopoulos, Jeffrey Sorensen, and Lucas Dixon. Civil rephrases of toxic texts with self-supervised transformers. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1442–1461. Association for Computational Linguistics.
  64. 64.Zhengzhe Liu, Xiaojuan Qi, and Philip H.S. Torr. Global texture enhancement for fake face detection in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  65. 65.Ed Pizzi, Giorgos Kordopatis-Zilos, Hiral Patel, Gheorghe Postelnicu, Sugosh Nagavara Ravindra, Akshay Gupta, Symeon Papadopoulos, Giorgos Tolias, and Matthijs Douze. The 2023 video similarity dataset and challenge. Computer Vision and Image Understanding, page 103997, 2024.
  66. 66.Wenhao Wang, Weipu Zhang, Yifan Sun, and Yi Yang. Bag of tricks and a strong baseline for image copy detection. arXiv preprint arXiv:2111.08004, 2021.
  67. 67.Zoe Papakipos, Giorgos Tolias, Tomas Jenicek, Ed Pizzi, Shuhei Yokoo, Wenhao Wang, Yifan Sun, Weipu Zhang, Yi Yang, Sanjay Addicam, et al. Results and findings of the 2021 image similarity challenge. In NeurIPS 2021 Competitions and Demonstrations Track, pages 1–12. PMLR, 2022.

Citation

MLA
Wang, W., and Y. Yang. “VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”. arXiv, 2024, http://arxiv.org/abs/2403.06098v4.
APA
Wang, W., & Yang, Y. (2024). VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models. arXiv. http://arxiv.org/abs/2403.06098v4
Chicago
Wang, W., and Y. Yang. 2024. “VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”. arXiv. http://arxiv.org/abs/2403.06098v4.
Harvard
Wang, W. and Yang, Y. (2024) “VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.06098v4.
Vancouver
1. Wang W, Yang Y (2024) VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models. arXiv

BibTeX

@article{wang2024vidprom,
  title = {VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models},
  author = {Wang, Wenhao and Yang, Yi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.06098v4},
  eprint = {2403.06098}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors