Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal datasets

A multimodal dataset is a curated collection of data used in machine learning that incorporates two or more distinct types or modalities of information, such as text, images, audio, video, or sensor readings. These datasets typically feature aligned or paired data points, such as text descriptions paired with corresponding images or video clips, which enable algorithms to learn associations and joint representations across different perceptual formats. By capturing relationships between diverse media forms, multimodal datasets provide the essential training and evaluation foundation for modern artificial intelligence systems performing tasks such as cross-modal search, visual question answering, and text-driven image or video generation.

2 items

CelebV-Text: A Large-Scale Facial Text-Video Dataset

CelebV-Text: A Large-Scale Facial Text-Video Dataset

Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy, Weidong Cai, Wayne Wu

Why you should read this

Presents a large-scale dataset of 70,000 in-the-wild facial video clips paired with 1.4 million detailed static and dynamic text descriptions to advance and standardize face-centric text-to-video generation.

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to-video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The superiority of CelebV-Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV-Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available^1.

Added

2026-09-26

DataComp: In search of the next generation of multimodal datasets

DataComp: In search of the next generation of multimodal datasets

Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander J. Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, Ludwig Schmidt

OrganizationsAllen Institute for AIAppleColumbia UniversityForschungszentrum JülichGoogleGraz University of TechnologyLAIONSnorkel AITel Aviv UniversityThe Hebrew University of JerusalemUniversity of Illinois Urbana-ChampaignUniversity of Texas at AustinUniversity of Washington

Why you should read this

Introduces DataComp, a standardized benchmark centered on 12.8 billion candidate image-text pairs that allows researchers to rigorously evaluate multimodal dataset filtering methods across multiple compute scales and produce CLIP models that surpass OpenAI's original zero-shot ImageNet accuracy.

Multimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the ML ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. In particular, our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release DataComp and all accompanying code at this http URL.

Added

2026-09-26