VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
Wenhao WangYi Yang
Introduces VidProM, the first large-scale prompt-gallery dataset containing 1.67 million real-user text-to-video prompts paired with 6.69 million synthesized videos from four diffusion models to advance research in video prompt engineering, efficient generation, and synthetic video detection.
Recent advances in artificial intelligence have enabled diffusion models to generate realistic video content from written instructions. However, these text-to-video systems depend heavily on the user prompts that guide them, and researchers previously lacked large-scale, real-world prompt datasets to study how users interact with these tools or to benchmark video generation performance effectively.
The article addresses this gap by introducing VidProM, the first large-scale prompt-gallery dataset curated specifically for text-to-video diffusion models. The primary objective is to demonstrate how video prompts fundamentally differ from image prompts and to provide a comprehensive public resource for evaluating models, improving generation efficiency, and developing safety and copyright safeguards.
To construct the dataset, the authors collected 1.67 million unique prompts submitted by real users on public Discord channels between July 2023 and February 2024. Using these prompts, they generated 6.69 million videos across four text-to-video diffusion systems (Pika, VideoCrafter2, Text2Video-Zero, and ModelScope), consuming over 50,000 GPU hours. Each prompt was processed with advanced text embeddings supporting up to 8,192 tokens and annotated with safety scores across six categories of potentially harmful or explicit content. The authors also filtered the collection to isolate approximately 1.04 million semantically unique prompts.
The analysis yielded several critical findings. First, video prompts differ sharply from image prompts: they are significantly longer (with roughly 60,000 prompts exceeding 70 words compared to only 15,000 in image benchmarks), incorporate temporal and dynamic descriptions, and focus heavily on human actions rather than static artistic styles. Second, existing fake-image detectors perform poorly when applied to generated video frames, with diffusion-specific detectors achieving around 49% accuracy (essentially random chance). Third, while direct replication of copyrighted training material by generative models occurs in only a small fraction of outputs (for instance, around 2% in open-source systems and even less in commercial tools), current copy-detection systems fail to reliably flag diffusion-based replications. Finally, the authors demonstrated that fine-tuning language models on VidProM enables effective automated prompt completion for video generation.
These findings indicate that generative video requires specialized engineering, evaluation frameworks, and safety solutions rather than direct adaptations from image tools. Organizations deploying or governing generative video face measurable risks around copyright infringement and deepfake detection, as current visual inspection tools do not generalize to video artifacts. Furthermore, because training data often relies on passive captions rather than user-style prompts, bridging this domain gap is necessary to improve commercial generation quality.
The authors recommend using VidProM to benchmark model performance against realistic user queries and explore prompt-retrieval techniques to reduce computing costs by adapting existing outputs rather than generating videos from scratch. They also suggest developing specialized, end-to-end fake video detectors and pairing automated copy filters with human verification to protect intellectual property. Users should note the dataset's primary limitations: the generated videos are relatively short (between 1.6 and 3.0 seconds) and derived from earlier open-source models rather than the latest frontier systems, meaning future updates will be required to reflect rapidly advancing video quality.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Introduces the foundational architecture for extending diffusion models to video synthesis, establishing the generative paradigm that VidProM analyzes and benchmarks.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). Demonstrates high-definition text-to-video generation using cascaded diffusion models, outlining the core text-conditioned synthesis techniques studied in large-scale prompt galleries.
- Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). Pioneers text-to-video generation by adapting pre-trained text-to-image models with temporal layers, illustrating the text-conditioned video models that require prompt dataset analysis.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Establishes high-resolution video synthesis using latent diffusion models, providing the architectural foundation for modern open-source text-to-video systems.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). Investigates the role of large-scale dataset curation in latent video diffusion models, providing essential context on how training data and prompt design influence video generation quality.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). Proposes a comprehensive multi-dimensional benchmark suite for video generative models, framing the evaluation protocols and quality dimensions relevant to prompt-gallery datasets.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). Extends text-to-video diffusion modeling using expert transformer architectures and detailed automated prompt pipelines to synthesize longer, highly coherent video sequences.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). Leverages real-world video dynamics and diffusion transformer architectures to unify visual generation, editing, and prompt-guided manipulation tasks.
