CelebV-Text: A Large-Scale Facial Text-Video Dataset
Jianhui YuHao ZhuLiming JiangChen Change LoyWeidong CaiWayne Wu
Presents a large-scale dataset of 70,000 in-the-wild facial video clips paired with 1.4 million detailed static and dynamic text descriptions to advance and standardize face-centric text-to-video generation.
Generating realistic human face videos directly from text prompts is a rapidly evolving area of artificial intelligence with extensive commercial, entertainment, and communication applications. However, existing text-to-video systems struggle with face generation, frequently yielding low-quality visuals, unnatural temporal artifacts, and poor alignment between the written instructions and the rendered video. A primary bottleneck has been the lack of large-scale facial video datasets that combine high-resolution clips with precise, highly relevant text descriptions capturing both static facial traits and dynamic expressions over time.
The article demonstrates the construction and effectiveness of CelebV-Text, a large-scale multimodal dataset designed to standardize and advance facial text-to-video generation. It systematically evaluates how detailed annotations spanning static appearances, fine facial marks, lighting conditions, dynamic actions, and emotions enhance the capability of generative models to produce faithful facial videos.
To build the dataset, the authors developed a semi-automated pipeline comprising curated data collection, hybrid automated and manual annotation, and template-based text generation. The final dataset encompasses 70,000 video clips totaling approximately 279 hours, with all videos maintaining a resolution of at least 512x512 pixels. Each clip is paired with 20 distinct text descriptions, generating 1,400,000 total descriptions. Static attributes cover 40 general appearance classes, five detailed facial features, and six lighting conditions, while dynamic attributes track 37 action types, eight basic emotions, and six lighting directions with exact start and end timestamps. Using these structured labels, grammatical parsing trees and vocabulary substitution generated diverse, natural language captions.
The article establishes several key findings. First, CelebV-Text provides substantially richer language diversity and descriptive detail than previous facial datasets; its average text description length of 67.15 words is more than double that of MM-Vox (28.39 words) and CelebV-HQ (31.06 words), incorporating 174 unique nouns and 96 verbs. Second, cross-modal retrieval experiments confirmed superior text-video relevance across appearance, emotion, and action categories compared to existing benchmarks. Third, when training generative models, a baseline architecture trained solely on CelebV-Text produced face videos with higher visual fidelity and closer prompt adherence than a major state-of-the-art general video model with roughly 100 times more parameters trained on 75 times more data. Fourth, introducing test-time text interpolation significantly stabilized dynamic attribute transitions and temporal coherence.
These findings indicate that domain-specific, densely annotated data is substantially more effective and cost-efficient for specialized generative tasks than simply scaling up uncurated, general-purpose models. Organizations seeking to deploy high-fidelity digital avatars or video synthesis tools can achieve superior visual compliance and temporal stability without the massive compute budgets typically required for multi-billion-parameter foundation models. Additionally, the structured benchmark offers a standardized metric framework to evaluate generative quality and relevance.
Moving forward, practitioners and research teams should utilize the publicly available dataset, annotations, and processing tools as a standardized benchmark for facial video synthesis. Technical teams developing dynamic video models should adopt text-interpolation techniques or similar cross-modal dynamic encoders to improve temporal continuity during expression changes. Further exploration is recommended to scale dataset diversity, adapt general foundation models to specialized facial domains, and advance text-driven three-dimensional facial synthesis.
Regarding limitations, real-world video collection introduces inherent distribution skews, such as frontal lighting representing 71% of samples and head movements comprising roughly 60% of dynamic actions. Generating dynamic temporal state changes also remains technically challenging, leading to noticeable quality drops in complex multi-action prompts compared to static descriptions. Because biometric data was excluded and dataset access will be governed by formal institutional legality checks, confidence in the dataset's utility, technical integrity, and ethical baseline is high.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). This foundational work establishes the paradigm of pairing multi-sentence text descriptions with open-domain video clips, providing the core dataset framing that CelebV-Text adapts for face-centric video generation.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). This paper establishes web-scraped video-text alignment paradigms and datasets (WebVid), framing the data requirements that CelebV-Text specializes for facial video-text pairs.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). This paper establishes cascaded diffusion models for high-definition text-to-video generation, highlighting the architectural demands and dataset dependencies addressed by CelebV-Text.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This work introduces foundational space-time diffusion architectures for text-conditioned video synthesis that underpin modern text-to-video modeling and evaluation benchmarks.
- Paper: Towards Accurate Generative Models of Video: A New Metric & Challenges, Thomas Unterthiner et al. (2018). This paper introduces Fréchet Video Distance (FVD), the primary metric used throughout generative video literature to benchmark the visual quality and temporal coherence of synthesized clips.
- Paper: VoxCeleb2: Deep Speaker Recognition, Joon Son Chung et al. (2018). This work introduces large-scale in-the-wild celebrity video curation and facial tracking pipelines that inform the data gathering methodology in facial video datasets.
- Paper: VGGFace2: A Dataset for Recognising Faces across Pose and Age, Qiong Cao et al. (2017). This study details large-scale in-the-wild facial data curation and attribute variation across pose and age, foundational for fine-grained facial dataset structuring.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). This work generalizes video generation benchmarking into comprehensive multidimensional metrics that evaluate fine-grained subject consistency and prompt alignment beyond single-score baselines.
- Paper: VideoBooth: Diffusion-based Video Generation with Image Prompts, Yuming Jiang et al. (2024). This paper extends text-driven video generation by incorporating personalized image prompts to achieve high subject fidelity across temporal frames without test-time fine-tuning.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). This study advances text-to-video generation by introducing a space-time U-Net architecture designed for globally coherent and realistic motion synthesis.
- Paper: Hierarchical Spatio-temporal Decoupling for Text-to- Video Generation, Zhiwu Qing et al. (2024). This work develops a hierarchical spatio-temporal decoupling framework to balance sharp spatial visual quality and dynamic motion in text-conditioned video synthesis.
- Paper: VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models, Wenhao Wang et al. (2024). This research analyzes real user text prompts at scale to study prompt-to-video dynamics and enhance text-to-video diffusion model training.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). This work applies 3D parametric modeling and motion adapters to synthesize realistic one-shot talking portrait videos driven by audio and visual signals.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). This paper proposes decoupled visual-motional tokenization to unify video-language pre-training across understanding and generation tasks.
