CelebV-Text: A Large-Scale Facial Text-Video Dataset

Jianhui YuHao ZhuLiming JiangChen Change LoyWeidong CaiWayne Wu

article2023CVPR127 citationsBest Paper Candidate

Presents a large-scale dataset of 70,000 in-the-wild facial video clips paired with 1.4 million detailed static and dynamic text descriptions to advance and standardize face-centric text-to-video generation.

Listen

Generating realistic human face videos directly from text prompts is a rapidly evolving area of artificial intelligence with extensive commercial, entertainment, and communication applications. However, existing text-to-video systems struggle with face generation, frequently yielding low-quality visuals, unnatural temporal artifacts, and poor alignment between the written instructions and the rendered video. A primary bottleneck has been the lack of large-scale facial video datasets that combine high-resolution clips with precise, highly relevant text descriptions capturing both static facial traits and dynamic expressions over time.

The article demonstrates the construction and effectiveness of CelebV-Text, a large-scale multimodal dataset designed to standardize and advance facial text-to-video generation. It systematically evaluates how detailed annotations spanning static appearances, fine facial marks, lighting conditions, dynamic actions, and emotions enhance the capability of generative models to produce faithful facial videos.

To build the dataset, the authors developed a semi-automated pipeline comprising curated data collection, hybrid automated and manual annotation, and template-based text generation. The final dataset encompasses 70,000 video clips totaling approximately 279 hours, with all videos maintaining a resolution of at least 512x512 pixels. Each clip is paired with 20 distinct text descriptions, generating 1,400,000 total descriptions. Static attributes cover 40 general appearance classes, five detailed facial features, and six lighting conditions, while dynamic attributes track 37 action types, eight basic emotions, and six lighting directions with exact start and end timestamps. Using these structured labels, grammatical parsing trees and vocabulary substitution generated diverse, natural language captions.

The article establishes several key findings. First, CelebV-Text provides substantially richer language diversity and descriptive detail than previous facial datasets; its average text description length of 67.15 words is more than double that of MM-Vox (28.39 words) and CelebV-HQ (31.06 words), incorporating 174 unique nouns and 96 verbs. Second, cross-modal retrieval experiments confirmed superior text-video relevance across appearance, emotion, and action categories compared to existing benchmarks. Third, when training generative models, a baseline architecture trained solely on CelebV-Text produced face videos with higher visual fidelity and closer prompt adherence than a major state-of-the-art general video model with roughly 100 times more parameters trained on 75 times more data. Fourth, introducing test-time text interpolation significantly stabilized dynamic attribute transitions and temporal coherence.

These findings indicate that domain-specific, densely annotated data is substantially more effective and cost-efficient for specialized generative tasks than simply scaling up uncurated, general-purpose models. Organizations seeking to deploy high-fidelity digital avatars or video synthesis tools can achieve superior visual compliance and temporal stability without the massive compute budgets typically required for multi-billion-parameter foundation models. Additionally, the structured benchmark offers a standardized metric framework to evaluate generative quality and relevance.

Moving forward, practitioners and research teams should utilize the publicly available dataset, annotations, and processing tools as a standardized benchmark for facial video synthesis. Technical teams developing dynamic video models should adopt text-interpolation techniques or similar cross-modal dynamic encoders to improve temporal continuity during expression changes. Further exploration is recommended to scale dataset diversity, adapt general foundation models to specialized facial domains, and advance text-driven three-dimensional facial synthesis.

Regarding limitations, real-world video collection introduces inherent distribution skews, such as frontal lighting representing 71% of samples and head movements comprising roughly 60% of dynamic actions. Generating dynamic temporal state changes also remains technically challenging, leading to noticeable quality drops in complex multi-action prompts compared to static descriptions. Because biometric data was excluded and dataset access will be governed by formal institutional legality checks, confidence in the dataset's utility, technical integrity, and ethical baseline is high.

arXiv: 2303.14717
Cover for CelebV-Text: A Large-Scale Facial Text-Video Dataset

Abstract

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to-video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The superiority of CelebV-Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV-Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available^1.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Dataset Construction
  • 3.1. Data Collection & Processing
  • 3.2. Data Annotation
  • 3.3. Semi-auto Text Generation
  • 4. Statistical Analysis of CelebV-Text
  • 4.1. Video Comparisons
  • 4.2. Text Comparisons
  • 4.3. Text-Video Relevance
  • 5. Experiment
  • 5.1. High-relevance Text-to-Video Generation
  • 5.2. Benchmark on Facial Text-to-Video Generation
  • 6. Discussion
  • References

Knowls

  1. Knowl 1 — CelebV-Text Dataset Specification and Taxonomy

    definition

    CelebV-Text is a large-scale, high-resolution facial text-video dataset designed for facial text-to-video generation tasks. It comprises 70,00070{,}000 in-the-wild facial video clips (totaling ∼279\sim 279 hours) paired with 1,400,0001{,}400{,}000 natural language descriptions (2020 descriptions per video clip). All video clips have a minimum resolution of 512×512512 \times 512, with 56.4%56.4\% having resolutions between 5122512^2 and 102421024^2, and 43.6%43.6\% exceeding 102421024^2.

    The dataset decouples video content into six attribute categories across static and dynamic domains:

    1. Static Attributes:
    • General Appearance: 40 discrete binary classes following CelebA definitions.
    • Detailed Appearance: 5 fine-grained classes (scar, mole, freckle, dimple, one-eyed) annotated with descriptive text specifying relative spatial locations on the face.
    • Light Conditions: 6 discrete classes capturing illumination brightness and correlated color temperature.
    1. Dynamic Attributes:
    • Action: 37 facial motion classes (such as head nodding, gazing, talking, squinting, blinking).
    • Emotion: 8 expression categories following AffectNet (neutral, anger, contempt, disgust, fear, happiness, sadness, surprise).
    • Light Direction: 6 directional illumination categories (e.g., front, back, left/right 45∘45^\circ, left/right 90∘90^\circ).

    All dynamic attributes are annotated with explicit start and end timestamps for continuous temporal localization.

  2. Knowl 2 — In-the-Wild Face Video Dataset Comparison

    data/table

    The scale, resolution, attribute coverage, and text annotation modality of CelebV-Text are benchmarked against existing in-the-wild facial video datasets:

    Dataset # Samples Resolution Duration Static Labels Dynamic Labels Timestamps Text Descriptions
    CelebV 5 256×256256 \times 256 2 hrs No No No None
    VoxCeleb2 150,480 224×224224 \times 224 2,442 hrs No No No None
    CelebV-HQ 35,666 512×512512 \times 512 68 hrs General App. Action Action only None
    MM-Vox 19,522 224×224224 \times 224 323 hrs General App. None No Automated
    CelebV-Text 70,000 ≥512×512\ge 512 \times 512 279 hrs General, Detail, Light Action, Emotion, Light Dir. All Dynamic Semi-automated (Auto + Manual)

    CelebV-Text provides higher resolution, densely timestamped dynamic actions/emotions, fine-grained detailed facial features, and paired natural language texts covering both static and dynamic factors.

  3. Knowl 3 — Semi-Automatic Template-Based Text Generation Pipeline

    model/method

    The semi-automatic text generation method produces diverse, grammatically natural, and attribute-aligned text descriptions for video clips by integrating automated parsing with human-written spatial annotations:

    1. Grammar Structure Extraction: Human annotators write ten sample descriptions per facial attribute. Syntactic parse trees from these descriptions are analyzed alongside external linguistic corpora to select the three most common grammatical structures for each of the six attribute categories, yielding a core library of 18 grammar templates.
    2. Probabilistic Context-Free Grammar (PCFG): Templates are parameterized using PCFG grammars to support combinatoric sentence generation.
    3. Manual-Text Injection: Un-discretized detailed appearance attributes (such as scars, moles, or freckles) with precise relative facial positions (written manually by annotators) are parsed and embedded directly into the structural templates.
    4. Lexical Diversification: Synonym replacement is performed on non-terminal lexical nodes using the Natural Language Toolkit (NLTK) to maximize vocabulary size and syntactic variety.
  4. Knowl 4 — Linguistic Diversity and POS Tag Distribution

    data/table

    The linguistic diversity of generated text descriptions is evaluated across part-of-speech (POS) tag counts and average word lengths per caption:

    Dataset # Verb # Adj. # Noun # Adv. Mean Word Length
    MM-Vox 5 20 38 0 28.39
    CelebV-HQ 10 24 50 6 31.06
    CelebV-Text 96 78 174 24 67.15

    CelebV-Text texts exhibit over double the average sentence length of prior datasets and incorporate significantly more unique verbs, adjectives, nouns, and adverbs, enabling fine-grained and temporally structured captioning.

  5. Knowl 5 — Multimodal Video-Text Retrieval Benchmark

    data/table

    Cross-modal relevance between video clips and text descriptions is evaluated using the Clip2Video temporal retrieval framework. Retrieval performance is reported via Recall at rank KK (R@K↑\text{R}@K \uparrow, in %\%), Median Rank (MdR↓\text{MdR} \downarrow), and Mean Rank (MnR↓\text{MnR} \downarrow) across both Text-to-Video (Text⇒Video\text{Text} \Rightarrow \text{Video}) and Video-to-Text (Video⇒Text\text{Video} \Rightarrow \text{Text}) tasks:

    Text ⇒\Rightarrow Video Video ⇒\Rightarrow Text
    Description Dataset R@1 R@5 R@10 MdR MnR R@1 R@5 R@10 MdR MnR
    (a) App. MM-Vox 1.5 9.0 15.7 52.0 68.8 2.0 9.2 14.6 43.0 57.8
    CelebV-HQ 5.9 19.2 29.7 27.0 52.2 7.2 20.7 32.4 27.0 46.9
    CelebV-Text 6.1 21.3 35.5 26.3 49.1 7.4 20.7 29.9 26.6 48.3
    (b) App.+Emo. CelebV-HQ 6.5 20.1 30.8 25.0 48.0 7.9 25.5 38.8 17.0 37.0
    CelebV-Text 6.6 23.4 37.1 26.0 47.6 8.1 27.2 34.7 18.2 38.3
    (c) App.+Emo.+Act. CelebV-Text 6.9 24.1 39.2 25.8 46.7 8.0 27.6 37.1 16.7 36.1

    Adding fine-grained dynamic actions and emotions incrementally improves retrieval scores, confirming the semantic relevance and alignment between the generated text annotations and the video content.

  6. Knowl 6 — Facial Text-to-Video Generation Benchmark

    data/table

    Facial text-to-video generation is evaluated on TFGAN, MMVID, and MMVID-interp using Fréchet Video Distance (FVD↓\text{FVD} \downarrow), Fréchet Inception Distance (FID↓\text{FID} \downarrow), and CLIP similarity (CLIPSIM↑\text{CLIPSIM} \uparrow). Mean values and standard errors over ten evaluation runs are reported:

    Condition Dataset / Method FVD (↓\downarrow) FID (↓\downarrow) CLIPSIM (↑\uparrow)
    General Appearance MM-Vox
    TFGAN 502.28±1.66502.28 \pm 1.66 760.24±16.01760.24 \pm 16.01 0.165±0.0220.165 \pm 0.022
    MMVID 65.79±1.8165.79 \pm 1.81 38.81±3.6638.81 \pm 3.66 0.170±0.0200.170 \pm 0.020
    CelebV-HQ
    TFGAN 428.04±1.76428.04 \pm 1.76 616.24±17.45616.24 \pm 17.45 0.168±0.0210.168 \pm 0.021
    MMVID 73.65±1.4373.65 \pm 1.43 63.86±3.6663.86 \pm 3.66 0.172±0.0190.172 \pm 0.019
    CelebV-Text
    TFGAN 403.04±1.34403.04 \pm 1.34 589.24±16.46589.24 \pm 16.46 0.177±0.0120.177 \pm 0.012
    MMVID 66.69±1.3566.69 \pm 1.35 58.70±4.6758.70 \pm 4.67 0.198±0.0140.198 \pm 0.014
    App. + Emotion CelebV-Text
    TFGAN 442.30±2.56442.30 \pm 2.56 623.17±18.88623.17 \pm 18.88 0.158±0.0240.158 \pm 0.024
    MMVID 82.78±1.4782.78 \pm 1.47 61.58±3.9961.58 \pm 3.99 0.176±0.0080.176 \pm 0.008
    MMVID-interp 72.87±1.2372.87 \pm 1.23 41.57±3.5641.57 \pm 3.56 0.182±0.0100.182 \pm 0.010
    App. + Action CelebV-Text
    TFGAN 571.34±4.54571.34 \pm 4.54 784.93±20.13784.93 \pm 20.13 0.154±0.0280.154 \pm 0.028
    MMVID 109.25±2.11109.25 \pm 2.11 82.55±4.3782.55 \pm 4.37 0.174±0.0190.174 \pm 0.019
    MMVID-interp 80.81±2.5580.81 \pm 2.55 70.88±4.7770.88 \pm 4.77 0.176±0.0200.176 \pm 0.020

    MMVID trained on CelebV-Text achieves superior CLIPSIM on general appearance text descriptions compared to prior datasets. When generating dynamic facial actions and emotions, MMVID-interp consistently improves visual quality (FID) and temporal consistency (FVD) over vanilla MMVID.

  7. Knowl 7 — MMVID-interp for Temporal Dynamic Facial Synthesis

    model/method

    MMVID-interp modifies the test-time text conditioning mechanism in discrete latent text-to-video generation (MMVID) to improve temporal coherence and attribute persistence during dynamic facial transitions (such as actions and emotional changes).

    Instead of conditioning frame sequence synthesis on step-function or abrupt sentence-level discrete embeddings, MMVID-interp interpolates text embedding representations continuously across adjacent dynamic intervals. This test-time interpolation smooths state transitions across temporal boundaries, stabilizes latent autoregressive code sampling, prevents the disappearance of static facial accessories (such as earrings or glasses) across temporal frames, and reduces FVD and FID on dynamic video generation.

Coverage note — Deliberately omitted qualitative comparison figures, background reviews of general video synthesis methods, and implementation details of off-the-shelf automated labeling classifiers found in the appendix.

References

  1. 1.Niki Aifanti, Christos Papachristou, and Anastasios Delopoulos. The mug facial expression database. In WIAMIS, 2010. 2
  2. 2.Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien. Unsupervised learning from narrated instruction videos. In CVPR, 2016. 1
  3. 3.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In ICCV, 2017. 2, 4
  4. 4.Max Bain, Arsha Nagrani, Gùl Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021. 1, 4
  5. 5.Yogesh Balaji, Martin Renqiang Min, Bing Bai, Rama Chellappa, and Hans Peter Graf. Conditional gan with discriminative filter generation for text-to-video synthesis. In IJCAI, 2019. 1, 2, 4, 6, 7, 8
  6. 6.Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple online and realtime tracking. In ICIP. IEEE, 2016. 3
  7. 7.Sergey Bezryadin, Pavel Bourov, and Dmitry Ilinih. Brightness calculation in digital image processing. In TDPF, 2007. 4
  8. 8.Steven Bird, Ewan Klein, and Edward Loper. Natural language processing with Python: analyzing text with the natural language toolkit. ” O’Reilly Media, Inc.”, 2009. 4
  9. 9.David Chen and William B Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011. 2, 4
  10. 10.Danqi Chen and Christopher D Manning. A fast and accurate dependency parser using neural networks. In EMNLP, 2014. 2, 4
  11. 11.Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. Voxceleb2: Deep speaker recognition. In INTERSPEECH, 2018. 1, 2, 3, 5
  12. 12.Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman. Quantifying generalization in reinforcement learning. In ICML, 2019. 2
  13. 13.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: collection, pipeline and challenges for epic-kitchens-100. In IJCV, 2022. 2
  14. 14.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019. 3
  15. 15.Kangle Deng, Tianyi Fei, Xin Huang, and Yuxin Peng. Ircgan: Introspective recurrent convolutional gan for text-tovideo generation. In IJCAI, pages 2216–2222, 2019. 2
  16. 16.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021. 2
  17. 17.Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097, 2021. 2, 6
  18. 18.Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi. Imagine this! scripts to compositions to videos. In ECCV, 2018. 1
  19. 19.Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, and Sergey Tulyakov. Show me what and tell me how: Video synthesis via multimodal conditioning. In CVPR, 2022. 2, 3, 4, 5, 6, 7, 8
  20. 20.William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. Flexible diffusion modeling of long videos. arXiv preprint arXiv:2205.11495, 2022. 2
  21. 21.Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Isabelle Hu, and Devi Parikh. Mugen: A playground for video-audio-text multimodal understanding and generation. arXiv preprint arXiv:2204.08058, 2022. 2, 4
  22. 22.Javier Hernandez-Andres, Raymond L Lee, and Javier Romero. Calculating correlated color temperatures across the entire gamut of daylight and skylight chromaticities. In Applied optics, 1999. 4
  23. 23.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. NeurIPS, 2017. 8
  24. 24.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2, 6, 7
  25. 25.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022. 2
  26. 26.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 2, 6, 7
  27. 27.Yaosi Hu, Chong Luo, and Zhenzhong Chen. Make it move: controllable image-to-video generation with text descriptions. In CVPR, 2022. 2, 4, 6
  28. 28.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, 2017. 2
  29. 29.Peter Kán and Hannes Kafumann. Deeplight: light source ´ estimation for augmented reality using deep learning. The Visual Computer, 2019. 4
  30. 30.Doyeon Kim, Donggyu Joo, and Junmo Kim. Tivgan: Text to image to video generation with step-by-step evolutionary generator. In IEEE Access, 2020. 3
  31. 31.Dan Klein and Christopher D Manning. Accurate unlexicalized parsing. In ACL, 2003. 4
  32. 32.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In ICCV, 2017. 2, 4
  33. 33.Dingquan Li, Tingting Jiang, and Ming Jiang. Quality assessment of in-the-wild videos. In ACM MM, pages 2351–2359, 2019. 5
  34. 34.Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin. Video generation from text. In Proceedings of the AAAI conference on artificial intelligence, 2018. 1
  35. 35.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 2
  36. 36.Wen Liu, Zhixin Piao, Jie Min, Wenhan Luo, Lin Ma, and Shenghua Gao. Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis. In ICCV, 2019. 2
  37. 37.Yue Liu, Xin Wang, Yitian Yuan, and Wenwu Zhu. Crossmodal dual learning for sentence-to-video generation. In ACM MM, 2019. 1
  38. 38.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In ICCV, 2015. 3, 4
  39. 39.Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, and Kevin Murphy. What’s cookin’? interpreting cooking videos using text, speech and vision. In ACL, 2015. 1
  40. 40.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In CVPR, 2019. 2, 4
  41. 41.Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K Roy-Chowdhury. Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In ICMR, 2018. 6
  42. 42.Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. In TIP, 2012. 5
  43. 43.Gaurav Mittal, Tanya Marwah, and Vineeth N Balasubramanian. Sync-draw: Automatic video generation using deep recurrent attentive architectures. In ACM MM, 2017. 1, 2
  44. 44.Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor. Affectnet: A database for facial expression, valence, and arousal computing in the wild. In TAC, 2017. 4
  45. 45.A. Nagrani, J. S. Chung, and A. Zisserman. Voxceleb: a large-scale speaker identification dataset. In INTERSPEECH, 2017. 1, 3
  46. 46.Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions. In ACM MM, 2017. 2
  47. 47.MarcAurelio Ranzato, Arthur Szlam, Joan Bruna, Michael Mathieu, Ronan Collobert, and Sumit Chopra. Video (language) modeling: a baseline for generative models of natural videos. arXiv preprint arXiv:1412.6604, 2014. 1
  48. 48.Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics: A large-scale video dataset for forgery detection in human faces. arXiv preprint arXiv:1803.09179, 2018. 1
  49. 49.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 2
  50. 50.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 2
  51. 51.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 2, 6, 7
  52. 52.David Stap, Maurits Bleeker, Sarah Ibrahimi, and Maartje ter Hoeve. Conditional image generation and manipulation for user-specified content. arXiv preprint arXiv:2005.04909, 2020. 2, 4
  53. 53.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 8
  54. 54.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. NeurIPS, 2017. 2
  55. 55.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description. arXiv preprint arXiv:2210.02399, 2022. 2, 6, 7
  56. 56.Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. Vatex: A large-scale, highquality multilingual dataset for video-and-language research. In ICCV, 2019. 4, 5, 6
  57. 57.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021. 2, 8
  58. 58.Chenfei Wu, Jian Liang, Xiaowei Hu, Zhe Gan, Jianfeng Wang, Lijuan Wang, Zicheng Liu, Yuejian Fang, and Nan Duan. Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis. In NeurIPS, 2022. 2
  59. 59.Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. N\” uwa: Visual synthesis pretraining for neural visual world creation. In ECCV, 2022. 2
  60. 60.Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. Reenactgan: Learning to reenact faces via boundary transfer. In ECCV, 2018. 1, 3, 5
  61. 61.Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. Tedigan: Text-guided diverse face image generation and manipulation. In CVPR, 2021. 3, 4
  62. 62.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016. 2, 4
  63. 63.Gilbert Youmans. Measuring lexical style and competence: The type-token vocabulary curve. Style, 1990. 6
  64. 64.Bowen Zhang, Hexiang Hu, and Fei Sha. Cross-modal and hierarchical modeling of video and text. In ECCV, 2018. 6
  65. 65.Luowei Zhou, Chenliang Xu, and Jason J Corso. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018. 2
  66. 66.Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. CelebVHQ: A large-scale video facial attributes dataset. In ECCV, 2022. 2, 3, 4, 5, 6, 7
  67. 67.Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Crosstask weakly supervised learning from instructional videos. In CVPR, 2019. 1, 4

Citation

MLA
Yu, J., et al. “CelebV-Text: A Large-Scale Facial Text-Video Dataset”. arXiv, 2023, http://arxiv.org/abs/2303.14717v1.
APA
Yu, J., Zhu, H., Jiang, L., Loy, C. C., Cai, W., & Wu, W. (2023). CelebV-Text: A Large-Scale Facial Text-Video Dataset. arXiv. http://arxiv.org/abs/2303.14717v1
Chicago
Yu, J., H. Zhu, L. Jiang, C. C. Loy, W. Cai, and W. Wu. 2023. “CelebV-Text: A Large-Scale Facial Text-Video Dataset”. arXiv. http://arxiv.org/abs/2303.14717v1.
Harvard
Yu, J. et al. (2023) “CelebV-Text: A Large-Scale Facial Text-Video Dataset”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.14717v1.
Vancouver
1. Yu J, Zhu H, Jiang L, Loy CC, Cai W, Wu W (2023) CelebV-Text: A Large-Scale Facial Text-Video Dataset. arXiv

BibTeX

@article{yu2023celebv,
  title = {CelebV-Text: A Large-Scale Facial Text-Video Dataset},
  author = {Yu, Jianhui and Zhu, Hao and Jiang, Liming and Loy, Chen Change and Cai, Weidong and Wu, Wayne},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.14717v1},
  eprint = {2303.14717}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE