DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak GadreGabriel IlharcoAlex FangJonathan HayaseGeorgios SmyrnisThao NguyenRyan MartenMitchell WortsmanDhruba GhoshJieyu Zhang
Introduces DataComp, a standardized benchmark centered on 12.8 billion candidate image-text pairs that allows researchers to rigorously evaluate multimodal dataset filtering methods across multiple compute scales and produce CLIP models that surpass OpenAI's original zero-shot ImageNet accuracy.
Recent breakthroughs in multimodal artificial intelligence rely heavily on massive web-scraped datasets, yet dataset curation has received significantly less rigorous research attention than model architectures or training algorithms. Many leading datasets remain proprietary, and existing public alternatives often suffer from unknown filtering impacts, toxic content, and inefficient scaling. To address these challenges, the article introduces DATACOMP, a standardized benchmarking testbed designed to foster systematic, data-centric research by holding model architectures and computational training budgets constant while evaluating different dataset curation strategies across 38 downstream visual and multimodal evaluation tasks.
The benchmark evaluates candidate pools spanning four orders of magnitude (from 12.8 million to 12.8 billion samples) and provides a curated starting reservoir called COMMONPOOL, which was harvested from Common Crawl with rigorous automated safety checks, face blurring, and deduplication against evaluation sets. Researchers can participate either by designing filtering techniques on COMMONPOOL or by bringing external data sources under the Bring Your Own Data track. Through over three hundred baseline experiments, the article demonstrates that dataset curation and quality matter far more than sheer volume.
The key finding shows that more stringently filtered subsets consistently outperform larger, uncurated datasets. In particular, combining image embedding clustering with multimodal similarity filtering created a new dataset, DATACOMP-1B, containing 1.4 billion samples. A standard Vision Transformer model trained from scratch on DATACOMP-1B achieved a 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI’s proprietary model by 3.7 percentage points and beating a model trained on LAION-2B by 6.1 percentage points while using the same or substantially less compute. This baseline delivered an approximate ninefold reduction in computational training costs relative to larger models trained on less curated pools.
These findings demonstrate that strategic data filtering offers organizations substantial cost savings, reduced computational resource requirements, and improved model performance without architectural modifications. Filtering ranking strategies proved remarkably consistent across compute scales and across different model architectures, enabling teams to prototype curation techniques cheaply at smaller scales before deploying them at massive scales. To maximize real-world deployment value, practitioners are recommended to prioritize rigorous data filtering workflows over brute-force web-scraping expansions.
Despite these advancements, users must remain cautious regarding residual data risks. While automated filters effectively strip explicit material and obfuscate faces, web-sourced data can still harbor demographic and socioeconomic biases, as seen in evaluations where lower-income categories underperformed. Further work is required to explore richer captioning supervision signals, expand into additional modalities such as video, and improve dataset balancing techniques without causing training divergence.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the contrastive language-image pre-training (CLIP) framework and zero-shot evaluation protocol that DataComp adopts as its standardized training and testing testbed.
- Paper: Reproducible Scaling Laws for Contrastive Language-Image Learning, Mehdi Cherti et al. (2022). Establishes the reproducible scaling laws and open-source OpenCLIP benchmarking infrastructure upon which DataComp's multi-scale compute tiers are directly constructed.
- Paper: LAION-5B: An open large-scale dataset for training next generation image-text models, Christoph Schuhmann et al. (2022). Demonstrates the construction and CLIP-filtering of multi-billion web-scraped image-text pairs from Common Crawl, providing the foundation for DataComp's candidate pool and baselines.
- Paper: LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, Christoph Schuhmann et al. (2021). Pioneers the open-access curation pipeline for extracting, filtering, and packaging hundreds of millions of web image-alt-text pairs that DataComp systematizes into a benchmark.
- Paper: Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts, Soravit Changpinyo et al. (2021). Explores how relaxing web data filters to capture broader vocabulary and long-tail concepts directly informs DataComp's focus on dataset curation strategies.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). Introduces synthetic captioning and dataset bootstrapping techniques for noisy web data that serve as key methodological reference points for DataComp's filtering tracks.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). Critically investigates the long-term limits and compute-scaling trade-offs of the web-data filtering paradigms formalized in DataComp.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). Applies large-scale multimodal pre-training and web data alignment principles to scale vision foundation models up to 6 billion parameters.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). Extends data curation and multimodal scaling methodologies to modern vision-language models combining high-quality filtered datasets with advanced test-time reasoning.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). Builds on multimodal dataset scaling insights to develop native, unified pre-training recipes balancing hundreds of billions of text and vision tokens.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). Advances beyond standard zero-shot vision-language evaluation suites to benchmark multi-discipline, expert-level reasoning in models trained on large multimodal datasets.
