DataComp: In search of the next generation of multimodal datasets

Samir Yitzhak GadreGabriel IlharcoAlex FangJonathan HayaseGeorgios SmyrnisThao NguyenRyan MartenMitchell WortsmanDhruba GhoshJieyu Zhang

article2023NeurIPS657 citations

Introduces DataComp, a standardized benchmark centered on 12.8 billion candidate image-text pairs that allows researchers to rigorously evaluate multimodal dataset filtering methods across multiple compute scales and produce CLIP models that surpass OpenAI's original zero-shot ImageNet accuracy.

Listen

Recent breakthroughs in multimodal artificial intelligence rely heavily on massive web-scraped datasets, yet dataset curation has received significantly less rigorous research attention than model architectures or training algorithms. Many leading datasets remain proprietary, and existing public alternatives often suffer from unknown filtering impacts, toxic content, and inefficient scaling. To address these challenges, the article introduces DATACOMP, a standardized benchmarking testbed designed to foster systematic, data-centric research by holding model architectures and computational training budgets constant while evaluating different dataset curation strategies across 38 downstream visual and multimodal evaluation tasks.

The benchmark evaluates candidate pools spanning four orders of magnitude (from 12.8 million to 12.8 billion samples) and provides a curated starting reservoir called COMMONPOOL, which was harvested from Common Crawl with rigorous automated safety checks, face blurring, and deduplication against evaluation sets. Researchers can participate either by designing filtering techniques on COMMONPOOL or by bringing external data sources under the Bring Your Own Data track. Through over three hundred baseline experiments, the article demonstrates that dataset curation and quality matter far more than sheer volume.

The key finding shows that more stringently filtered subsets consistently outperform larger, uncurated datasets. In particular, combining image embedding clustering with multimodal similarity filtering created a new dataset, DATACOMP-1B, containing 1.4 billion samples. A standard Vision Transformer model trained from scratch on DATACOMP-1B achieved a 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI’s proprietary model by 3.7 percentage points and beating a model trained on LAION-2B by 6.1 percentage points while using the same or substantially less compute. This baseline delivered an approximate ninefold reduction in computational training costs relative to larger models trained on less curated pools.

These findings demonstrate that strategic data filtering offers organizations substantial cost savings, reduced computational resource requirements, and improved model performance without architectural modifications. Filtering ranking strategies proved remarkably consistent across compute scales and across different model architectures, enabling teams to prototype curation techniques cheaply at smaller scales before deploying them at massive scales. To maximize real-world deployment value, practitioners are recommended to prioritize rigorous data filtering workflows over brute-force web-scraping expansions.

Despite these advancements, users must remain cautious regarding residual data risks. While automated filters effectively strip explicit material and obfuscate faces, web-sourced data can still harbor demographic and socioeconomic biases, as seen in evaluations where lower-income categories underperformed. Further work is required to explore richer captioning supervision signals, expand into additional modalities such as video, and improve dataset balancing techniques without causing training divergence.

Cover for DataComp: In search of the next generation of multimodal datasets

Abstract

Multimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the ML ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. In particular, our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release DataComp and all accompanying code at this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The DataComp benchmark
  • 3.1 Competition design
  • 3.2 CommonPool generation, for the filtering track
  • 3.3 The bring your own data (BYOD) track
  • 3.4 Training
  • 3.5 Evaluation
  • 4 Baselines
  • 4.1 Filtering baselines
  • 4.2 BYOD baselines
  • 5 Results and discussion
  • 5.1 Building better datasets
  • 5.2 DataComp design analyses
  • 5.3 Evaluation trends
  • 6 Limitations and conclusion
  • References
  • A Benchmark rules
  • A.1 Filtering track rules
  • A.2 Bring your own data track: amendments
  • B Contributions
  • B.1 Candidate pool
  • B.2 Participant tooling
  • B.3 Baselines
  • B.4 Leadership and Advising
  • C Additional related work
  • D Parsing Common Crawl
  • E Not safe for work (NSFW) filtering
  • F Deduplication against evaluation sets
  • G Face blurring
  • H DataComp CommonPool creation pipeline
  • I CommonPool statistics
  • J Efficient training on data subsets
  • K Effect of duplicates in the training data
  • L Hyperparameter ablations
  • L.1 Batch size
  • L.2 Model architecture
  • L.3 Number of training steps
  • M Detector-based baselines
  • N Training details
  • O Evaluation details
  • O.1 Visual Question Answering
  • P Baseline details
  • P.1 Filtering track
  • P.2 BYOD track
  • P.2.1 Additional results
  • Q Fairness and biases
  • Q.1 Diversity
  • Q.2 Fairness
  • R Extra figures and tables
  • S Datasheet
  • S.1 Motivation
  • S.2 Composition
  • S.3 Collection Process
  • S.4 Preprocessing, Cleaning, and/or Labeling
  • S.5 Uses
  • S.6 Distribution
  • S.7 Maintenance

Knowls

  1. Knowl 1 — DataComp Benchmark Architecture and Multiscale Training Protocol

    experimental setup

    The DataComp benchmark reverses the conventional machine learning evaluation paradigm: rather than fixing the dataset and evaluating different model architectures or training algorithms, it holds the entire training pipeline, model architecture, and compute budget fixed, benchmarking different data curation and filtering strategies.

    Pretraining employs contrastive language-image pretraining (CLIP) trained from scratch with the InfoNCE contrastive loss over image representations g(xi)g(x_i) and text representations h(yi)h(y_i) in a batch of size BB:

    ℓ=12B∑i=1Blog⁡exp⁡(⟨g(xi),h(yi)⟩/τ)∑j=1Bexp⁡(⟨g(xi),h(yj)⟩/τ)+12B∑i=1Blog⁡exp⁡(⟨g(xi),h(yi)⟩/τ)∑j=1Bexp⁡(⟨g(xj),h(yi)⟩/τ)\ell = \frac{1}{2 B} \sum_{i=1}^{B} \log \frac{\exp(\langle g(x_i), h(y_i) \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle g(x_i), h(y_j) \rangle / \tau)} + \frac{1}{2 B} \sum_{i=1}^{B} \log \frac{\exp(\langle g(x_i), h(y_i) \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle g(x_j), h(y_i) \rangle / \tau)}

    where τ\tau is a learnable temperature parameter.

    The benchmark defines four standardized compute scales spanning four orders of magnitude:

    1. Small: Candidate pool of 12.8×10612.8 \times 10^6 samples, 12.8×10612.8 \times 10^6 total training samples seen, Vision Transformer (ViT) architecture ViT-B/32, total compute 9.5×10169.5 \times 10^{16} multiply-accumulate operations (MACs), learning rate 5×10−45 \times 10^{-4}, AdamW β2=0.98\beta_2 = 0.98, warmup 500 steps, batch size 4096.
    2. Medium: Candidate pool of 128×106128 \times 10^6 samples, 128×106128 \times 10^6 total samples seen, ViT-B/32, total compute 9.5×10179.5 \times 10^{17} MACs, learning rate 5×10−45 \times 10^{-4}, AdamW β2=0.98\beta_2 = 0.98, warmup 500 steps, batch size 4096.
    3. Large: Candidate pool of 1.28×1091.28 \times 10^9 samples, 1.28×1091.28 \times 10^9 total samples seen, ViT-B/16, total compute 2.6×10192.6 \times 10^{19} MACs, learning rate 5×10−45 \times 10^{-4}, AdamW β2=0.98\beta_2 = 0.98, warmup 500 steps, batch size 8192.
    4. XLarge: Candidate pool of 12.8×10912.8 \times 10^9 samples, 12.8×10912.8 \times 10^9 total samples seen, ViT-L/14, total compute 1.1×10211.1 \times 10^{21} MACs, learning rate 1×10−31 \times 10^{-3}, AdamW β2=0.95\beta_2 = 0.95, warmup 10,000 steps, batch size 90,112.

    In the filtering track, participants must select a subset of the candidate pool at the chosen scale. When a participant selects a subset smaller than the fixed sample budget, training loops over that subset for multiple epochs until the prescribed sample budget is exhausted.

  2. Knowl 2 — CommonPool Construction and Multi-Stage Safety Curation Funnel

    model/method

    CommonPool is a public multimodal candidate pool containing 12.8 billion image-text pairs harvested from Common Crawl snapshots spanning 2014 through 2022. Its creation follows a sequential data pipeline:

    1. Extraction and Download: Image URLs and non-empty alt-text captions are parsed from Common Crawl metadata archives (WAT files) using Apache Spark, yielding approximately 88 billion deduplicated URL-text pairs. Attempting downloads on a 40 billion sample subset yields approximately 16.8 billion downloaded image-text pairs, where images are resized such that the maximum dimension does not exceed 512 pixels.
    2. Text Safety Filtering: Captions are evaluated using the multilingual Detoxify model (XLM-RoBERTa variant). A sample is discarded if any of its scores across seven categories (toxicity, severe toxicity, obscene, identity attack, insult, threat, sexually explicit) exceeds a threshold of 0.1.
    3. Visual NSFW Filtering: Visual content is filtered using a 4-layer multilayer perceptron classifier trained on CLIP ViT-L/14 image embeddings using the LAION-5B NSFW dataset. Using a classification threshold of 0.1, the filter achieves 97.4% accuracy on held-out test data. Visual and text filtering collectively remove ~19% of downloads, reducing the pool to ~13.6 billion samples.
    4. Evaluation Set Deduplication: An image copy detection model removes approximately 0.5 billion samples (~3% of the downloads) that are near-duplicates of images in the 38 downstream evaluation datasets, leaving ~13.1 billion viable pairs.
    5. Subsampling and Privacy: A uniform random sample of 12.8 billion image-text pairs forms the xlarge CommonPool. Telescoping random subsets of size 1.28B, 128M, and 12.8M form the large, medium, and small CommonPool datasets, respectively. The download tooling automatically detects and blurs human faces using the SCRFD-10G face detector at a confidence threshold of 0.3.
  3. Knowl 3 — DataComp-1B State-of-the-Art Zero-Shot Multimodal Performance

    data/table

    DataComp-1B is a 1.4 billion sample multimodal dataset created by taking the intersection of ImageNet-1K feature clustering (image-based filtering) and the top 30% CLIP ViT-L/14 cosine similarity score filtering from the 12.8B CommonPool candidate reservoir. Training a CLIP ViT-L/14 model from scratch on DataComp-1B achieves 79.2% zero-shot top-1 accuracy on ImageNet-1K, surpassing OpenAI's CLIP ViT-L/14 by 3.7 percentage points and LAION-2B ViT-L/14 by 6.1 percentage points under identical compute budgets (1.1×10211.1 \times 10^{21} MACs / 13B samples seen).

    Dataset Dataset Size Samples Seen Architecture Train Compute (MACs) ImageNet Top-1 Accuracy (%)
    OpenAI WIT 0.4B 13B ViT-L/14 1.1×10211.1 \times 10^{21} 75.5
    LAION-400M 0.4B 13B ViT-L/14 1.1×10211.1 \times 10^{21} 72.8
    LAION-2B 2.3B 13B ViT-L/14 1.1×10211.1 \times 10^{21} 73.1
    LAION-2B 2.3B 34B ViT-H/14 6.5×10216.5 \times 10^{21} 78.0
    LAION-2B 2.3B 34B ViT-g/14 9.9×10219.9 \times 10^{21} 78.5
    DataComp-1B 1.4B 13B ViT-L/14 1.1×10211.1 \times 10^{21} 79.2

    DataComp-1B achieves higher zero-shot accuracy than a much larger ViT-g/14 model pre-trained on LAION-2B for nearly 3×3\times more samples seen, representing an overall 9×9\times compute reduction.

  4. Knowl 4 — Generalization Advantage of Stringent Filtering under Fixed Compute Budgets

    empirical result

    When holding the total training compute (number of samples seen) constant, training for multiple epochs over smaller, stringently filtered datasets yields significantly higher zero-shot classification accuracy across diverse evaluation sets than single-epoch training on the entire unfiltered dataset.

    At the xlarge scale (12.8B samples seen during pretraining):

    1. No Filtering (12.8B unique samples, 1 epoch): 72.3% ImageNet zero-shot accuracy, 0.621 average accuracy across 38 tasks.
    2. LAION-2B Style Filtering (1.3B unique samples, ~10 epochs): 75.5% ImageNet accuracy, 0.636 38-task average.
    3. CLIP ViT-L/14 Score Top 30% Filtering (3.8B unique samples, ~3.4 epochs): 76.4% ImageNet accuracy, 0.650 38-task average.
    4. ImageNet-1K Image Clustering ∩\cap CLIP ViT-L/14 Score Top 30% (1.4B unique samples, ~9 epochs): 79.2% ImageNet accuracy, 0.663 38-task average.

    Conversely, taking uncurated random subsets of decreasing size (e.g., 75%, 50%, 25%, 10%, 1%) consistently and monotonically degrades downstream zero-shot accuracy compared to the full pool (at small scale: 12.8M random achieves 0.133 average score, while a 1% random subset of 128K achieves 0.078). This confirms that repeated exposure to curated, high-quality image-text alignments drives generalization gains rather than raw dataset size or simple repetition.

  5. Knowl 5 — Image-Based Dataset Filtering via Feature Space Clustering

    algorithm

    Image-based filtering identifies samples in an uncurated web pool whose visual features semantically match a clean reference dataset (such as ImageNet-1K or ImageNet-21K) without relying on text overlap.

    Input: Candidate pool image-text pairs (xj,yj)j=1M(x_j, y_j)_{j=1}^M, clean reference training images (zi)i=1N(z_i)_{i=1}^N, CLIP image encoder gg, number of clusters K=105K = 10^5, kk-means iterations T=20T = 20.
    Output: Filtered candidate subset index set JfilteredJ_{filtered}.
    Filter candidate pool by basic text criteria (retain only fasttext English with text length ≥2\ge 2 words and ≥6\ge 6 characters), yielding MM remaining items
    Extract candidate image embeddings ej=g(xj)∈Rde_j = g(x_j) \in \mathbb{R}^d for j∈{1,…,M}j \in \{1, \dots, M\}
    Run kk-means clustering on candidate embeddings for TT iterations to obtain KK cluster centroids c1,…,cK∈Rdc_1, \dots, c_K \in \mathbb{R}^d
    Define cluster assignment function I(v)=arg⁡max⁡1≤k≤K⟨v,ck⟩I(v) = \arg\max_{1 \le k \le K} \langle v, c_k \rangle
    Extract reference image embeddings fi=g(zi)f_i = g(z_i) for i∈{1,…,N}i \in \{1, \dots, N\}
    Compute set of active reference cluster indices S={I(fi)∣1≤i≤N}S = \{I(f_i) \mid 1 \le i \le N\}
    Initialize Jfiltered←∅J_{filtered} \leftarrow \emptyset
    for j=1j = 1 to MM do
        if I(ej)∈SI(e_j) \in S then
            Jfiltered←Jfiltered∪{j}J_{filtered} \leftarrow J_{filtered} \cup \{j\}
    return JfilteredJ_{filtered}

    When applied with OpenAI CLIP ViT-L/14 embeddings and ImageNet-1K as the reference set, this algorithm isolates visual clusters relevant to downstream visual recognition benchmarks, reducing the candidate pool to high-yield clusters.

  6. Knowl 6 — Cross-Scale and Architectural Consistency of Data Filtering Strategies

    empirical result

    The relative ranking of data filtering and curation strategies is strongly preserved across compute scales, batch sizes, and model architectures:

    1. Cross-Scale Rank Correlation: The Spearman rank correlation of downstream performance between filtering methods evaluated at different compute scales is consistently high:
      • Small scale (1.28×1071.28 \times 10^7 seen) vs. Medium scale (1.28×1081.28 \times 10^8 seen): 0.8950.895 for ImageNet accuracy, 0.8540.854 for the 38-task average.
      • Small scale vs. Large scale (1.28×1091.28 \times 10^9 seen): 0.8110.811 for ImageNet accuracy, 0.7080.708 for the 38-task average.
      • Medium scale vs. Large scale: 0.8470.847 for ImageNet accuracy, 0.8760.876 for the 38-task average.
    2. Batch Size Invariance: Doubling batch size at the medium scale (from 4096 to 8192) yields a rank correlation of 0.960.96 on ImageNet accuracy and 0.980.98 on the 38-task average metric, with rankings shifting by at most one position.
    3. Architecture Invariance: Replacing the Vision Transformer (ViT-B/32) with a compute-matched ConvNeXt model at the medium scale yields a rank correlation of 1.01.0 on ImageNet accuracy and 0.870.87 on the 38-task average.
    4. Training Steps Invariance: Extending pretraining duration by 10×10\times steps at the small scale maintains a positive performance correlation with standard compute configurations.

    These properties allow researchers to evaluate novel dataset curation algorithms at small and medium scales (4 to 40 GPU hours on an A100) with high predictive confidence for large-scale pretraining.

  7. Knowl 7 — Evaluation Set Deduplication via Contrastive Copy Detection

    model/method

    To prevent benchmark contamination and train-test data leakage, CommonPool incorporates a near-duplicate image deduplication step against all images in the 38 downstream evaluation datasets using the self-supervised copy detection descriptor model by Yokoo (first-place solution in the Facebook AI Image Similarity Challenge, ISC).

    An image from the candidate pool is flagged and removed if the cosine similarity between its descriptor feature vector and any evaluation image feature vector exceeds 0.6041690.604169. This threshold removes approximately 2.8% to 3.0% (~0.5 billion) of downloaded images.

    In controlled benchmarks against realistic image transformations (JPEG compression, horizontal flips, rotations, aspect ratio modifications, and grayscaling) under a 4:1 distractor-to-reference ratio:

    • The ISC copy detection model achieves a precision of 0.90.9 and a recall of 0.80.8 at the 0.6041690.604169 threshold.
    • Standard OpenAI CLIP ViT-B/32 features at a similarity threshold of 0.960.96 achieve the same 0.90.9 precision but suffer a recall of only 0.020.02.
    • To match the recall of the ISC model, CLIP ViT-L/14 thresholding falsely removes more than 2×2\times the amount of non-duplicate training data.
  8. Knowl 8 — DataComp Evaluation Benchmark Suite and Metric Alignments

    experimental setup

    The DataComp evaluation suite comprises 38 zero-shot vision-language downstream tasks spanning diverse distributions and capabilities:

    1. Task Breakdown:

      • 22 Standard CLIP Benchmarks: Caltech-101, CIFAR-10, CIFAR-100, Country211, DTD, EuroSAT, FGVC Aircraft, Food-101, GTSRB, ImageNet-1K, MNIST, Oxford Flowers-102, Oxford-IIIT Pet, Pascal VOC 2007, Rendered SST2, RESISC45, Stanford Cars, STL-10, SUN-397, SVHN.
      • 6 ImageNet Distribution Shifts: ImageNet-Sketch, ImageNet-V2, ImageNet-A, ImageNet-O, ImageNet-R, ObjectNet.
      • 13 Visual Task Adaptation Benchmark (VTAB) Tasks: Counting (CLEVR Counts), distance prediction (CLEVR Distance, KITTI Distance), pathology (PatchCamelyon, Camelyon17), satellite recognition, etc.
      • 3 In-the-Wild Robustness Tasks (WILDS): iWildCam (macro F1 score), Camelyon17, FMoW (worst-region accuracy).
      • 2 Socioeconomic and Geographic Diversity Tasks: Dollar Street (worst-income top-5 accuracy) and GeoDE (worst-region accuracy across 6 global regions).
      • 3 Retrieval and Association Tasks: Flickr30k (average image/text R@1), MSCOCO (average image/text R@1), and WinoGAViL (commonsense association Jaccard score).
    2. Protocol Validation: Across all baselines, zero-shot ImageNet top-1 accuracy achieves a linear correlation of 0.990.99 with average performance across the 38 tasks. Zero-shot ImageNet accuracy also exhibits a Spearman rank correlation of 0.990.99 (CommonPool) and 1.01.0 (BYOD) with linear probe fine-tuning performance (40 epochs, learning rate 10−310^{-3}), demonstrating that zero-shot evaluation accurately reflects linear probing representations.

  9. Knowl 9 — Multi-Source Data Blending in the Bring Your Own Data Track

    empirical result

    In the Bring Your Own Data (BYOD) track, augmenting CommonPool with external publicly available image-text datasets improves downstream zero-shot accuracy over single-source training:

    1. Large Scale (1.28B Seen Samples): Combining CLIP-filtered CommonPool with four external datasets (CC12M, YFCC15M, RedCaps, Shutterstock photos) totaling 109 million unique samples increases ImageNet top-1 zero-shot accuracy from 60.9% (single-pass CommonPool baseline) to 63.5% when the external sources are upsampled 6×6\times, and to 62.4% when a larger suite of external datasets is upsampled 8×8\times.
    2. XLarge Scale (12.8B Seen Samples): Combining CommonPool with the four external datasets upsampled 6×6\times achieves 77.6% ImageNet accuracy and 0.649 average accuracy over 38 tasks, outperforming LAION-2B (75.7% ImageNet, 0.621 average).
    3. Domain Balancing: Blending diverse web domains mitigates domain-specific vocabulary deficits, but excessive upsampling beyond 10×10\times to 18×18\times causes marginal returns or performance regression due to duplicate representation saturation.
  10. Knowl 10 — Failure of Naive Class and Spatial Balancing in Multimodal Contrastive Pretraining

    limitation

    Applying supervised object detection heuristics to enforce class, spatial, or count balance across pretraining datasets severely impairs multimodal contrastive learning.

    When annotating the 128M medium CommonPool with the Detic object detector across 1,203 LVIS categories (confidence threshold >0.5> 0.5):

    • No Filtering Baseline: 17.6% ImageNet accuracy, 0.258 38-task average.
    • Object Exists Filter (retaining images with ≥1\ge 1 LVIS detection): 18.1% ImageNet accuracy, 0.263 average.
    • Balance by Class (equal sample distribution across 1,204 class buckets): 3.8% ImageNet accuracy, 0.141 average.
    • Balance by Spatial Position (equal distribution across a 5×55 \times 5 image grid): 4.0% ImageNet accuracy, 0.148 average.
    • Balance by Object Count (equal distribution across 0 to >10>10 detections): 12.7% ImageNet accuracy, 0.221 average.
    • Class Balancing on Curated Data: Applying class balancing to the high-performing Image-based ∩\cap CLIP score baseline collapses ImageNet accuracy from 29.7% to 3.4%.

    This collapse occurs because enforcing class and spatial uniformity on long-tailed web distributions requires massive oversampling of rare categories, which repeatedly injects duplicate or near-identical image-text pairs into the same batch, causing the InfoNCE contrastive denominator to contrast positive pairs against identical negatives, leading to loss divergence.

  11. Knowl 11 — Demographic Biases, Face Obfuscation, and Harmful Misclassification Risks

    limitation

    Investigation of privacy safeguards and demographic fairness in models pre-trained on web-scraped data reveals significant ethical limitations:

    1. Face Blurring Effect on Downstream Utility: Detecting and blurring human faces with SCRFD-10G has a negligible effect on general downstream vision tasks. At the medium scale, models trained on blurred versus non-blurred data achieve 28.2% vs. 28.7% ImageNet accuracy and 0.298 vs. 0.301 average accuracy on the 38 evaluation tasks.
    2. Demographic Classification from Obfuscated Data: Despite face blurring, models still classify demographic attributes on FairFace and UTKFace significantly above chance (e.g., FairFace race accuracy of 86.4%, gender accuracy of 91.7%), likely relying on context, clothing, hair, or skin tone.
    3. Subgroup Disparities: On FairFace, the filtering track model misclassifies Black, Southeast Asian, and East Asian males as females at rates of 20.7%, 17.0%, and 19.3% respectively.
    4. Harmful Crime Stereotyping: In zero-shot association tests combining demographic prompts with crime labels ("thief", "criminal", "suspicious person") versus non-human terms ("animal", "chimpanzee", "gorilla"):
      • Misclassification as non-human is low (≤0.1%\le 0.1\% across all subgroups).
      • Misclassification into criminal categories occurs at alarmingly high rates across both datasets (e.g., up to 24.3% of White individuals on FairFace and up to 35.3% of Southeast Asian individuals in the BYOD track), highlighting the inherent societal biases present in web-scraped image-text datasets.

Coverage note — Omitted materials include exhaustive individual per-task accuracy tables across all 38 datasets (summarized via representative aggregate benchmarks and ImageNet/38-task metrics), low-level distributed infrastructure engineering scripts, and standard dataset documentation boilerplate from the appendix datasheet.

References

  1. 1.cc2dataset. https://github.com/rom1504/cc2dataset.
  2. 2.CLD3. https://github.com/google/cld3.
  3. 3.Common Crawl. https://commoncrawl.org.
  4. 4.dataset2metadata. https://github.com/mlfoundations/dataset2metadata.
  5. 5.img2dataset. https://github.com/rom1504/img2dataset.
  6. 6.Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023. https://arxiv.org/abs/2303.09540.
  7. 7.Pankaj K. Agarwal, Sariel Har-Peled, and Kasturi R. Varadarajan. Approximating extent measures of points. Journal of the ACM (JACM), 2004. https://doi.org/10.1145/1008731.1008736.
  8. 8.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://openreview.net/forum?id=EbMuimAbPbs.
  9. 9.Abhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, and Sunita Sarawagi. Learning from rules generalizing labeled exemplars. In International Conference on Learning Representations (ICLR), 2020. https://openreview.net/forum?id=SkeuexBtDr.
  10. 10.Stephen H Bach, Daniel Rodriguez, Yintao Liu, Chong Luo, Haidong Shao, Cassandra Xia, Souvik Sen, Alex Ratner, Braden Hancock, Houman Alborzi, Rahul Kuchhal, Christopher Ré, and Rob Malkin. Snorkel drybell: A case study in deploying weak supervision at industrial scale. In Special Interest Group on Management of Data (SIGMOD), 2019. https://arxiv.org/abs/1812.00417.
  11. 11.Olivier Bachem, Mario Lucic, and Andreas Krause. Coresets for nonparametric estimation - the case of dp-means. In International Conference on Machine Learning (ICML), 2015. https://proceedings.mlr.press/v37/bachem15.html.
  12. 12.Peter Bandi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, et al. From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge. IEEE Transactions on Medical Imaging, 2018. https://pubmed.ncbi.nlm.nih.gov/30716025/.
  13. 13.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (eds.), Advances in Neural Information Processing Systems (NeurIPS), volume 32. Curran Associates, Inc., 2019. https://proceedings.neurips.cc/paper/2019/file/97af07a14cacba681feacf3012730892-Paper.pdf.
  14. 14.Sara Beery, Elijah Cole, and Arvi Gjoka. The iwildcam 2020 competition dataset, 2020. https://arxiv.org/abs/2004.10340.
  15. 15.Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misogyny, pornography, and malignant stereotypes, 2021. https://arxiv.org/abs/2110.01963.
  16. 16.Vighnesh Birodkar, Hossein Mobahi, and Samy Bengio. Semantic redundancies in image-classification datasets: The 10% you don’t need. arXiv preprint arXiv:1901.11409, 2019. https://arxiv.org/abs/1901.11409.
  17. 17.Yonatan Bitton, Nitzan Bitton Guetta, Ron Yosef, Yuval Elovici, Mohit Bansal, Gabriel Stanovsky, and Roy Schwartz. WinoGAViL: Gamified association benchmark to challenge vision-and-language models, 2022. https://arxiv.org/abs/2207.12576.
  18. 18.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision (ECCV), 2014. https://link.springer.com/chapter/10.1007/978-3-319-10599-4_29.
  19. 19.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems (NeurIPS), 2020. https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  20. 20.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022.
  21. 21.Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger. Broken neural scaling laws. International Conference on Learning Representations (ICLR), 2023. https://arxiv.org/abs/2210.14891.
  22. 22.Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. Poisoning web-scale training datasets is practical, 2023. https://arxiv.org/abs/2302.10149.
  23. 23.Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, and Felix Hill. Data distributional properties drive emergent in-context learning in transformers. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://arxiv.org/abs/2205.05055.
  24. 24.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. https://arxiv.org/abs/2102.08981.
  25. 25.Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. Pali: A jointly-scaled multilingual language-image model. In International Conference on Learning Representations (ICLR), 2022. https://arxiv.org/abs/2209.06794.
  26. 26.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO captions: Data collection and evaluation server, 2015. https://arxiv.org/abs/1504.00325.
  27. 27.Gong Cheng, Junwei Han, and Xiaoqiang Lu. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the Institute of Electrical and Electronics Engineers (IEEE), 2017. https://ieeexplore.ieee.org/abstract/document/7891544.
  28. 28.Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning, 2022. https://arxiv.org/abs/2212.07143.
  29. 29.Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. https://arxiv.org/abs/1711.07846.
  30. 30.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2014. https://openaccess.thecvf.com/content_cvpr_2014/html/Cimpoi_Describing_Textures_in_2014_CVPR_paper.html.
  31. 31.Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2011. https://proceedings.mlr.press/v15/coates11a.html.
  32. 32.Michael B. Cohen, Cameron Musco, and Christopher Musco. Input sparsity time low-rank approximation via ridge leverage score sampling. In ACM-SIAM Symposium on Discrete Algorithms, 2017. https://dl.acm.org/doi/10.5555/3039686.3039801.
  33. 33.C Coleman, C Yeh, S Mussmann, B Mirzasoleiman, P Bailis, P Liang, J Leskovec, and M Zaharia. Selection via proxy: Efficient data selection for deep learning. In International Conference on Learning Representations (ICLR), 2020. https://arxiv.org/abs/1906.11829.
  34. 34.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Annual Meeting of the Association for Computational Linguistics (ACL), 2019. https://arxiv.org/abs/1911.02116.
  35. 35.R Dennis Cook. Detection of influential observation in linear. Technometrics, 19(1):15–18, 1977.
  36. 36.Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan. Tao: A large-scale benchmark for tracking any object. In European Conference on Computer Vision (ECCV), 2020. https://arxiv.org/abs/2005.10356.
  37. 37.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition (CVPR), 2009. https://ieeexplore.ieee.org/abstract/document/5206848.
  38. 38.Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson. Redcaps: Web-curated image-text data created by the people, for the people, 2021. https://arxiv.org/abs/2111.11431.
  39. 39.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021. https://openreview.net/forum?id=YicbFdNTTy.
  40. 40.Matthijs Douze, Giorgos Tolias, Ed Pizzi, Zoë Papakipos, Lowik Chanussot, Filip Radenovic, Tomas Jenicek, Maxim Maximov, Laura Leal-Taixé, Ismail Elezi, Ondrej Chum, and Cristian Canton-Ferrer. The 2021 image similarity dataset and challenge, 2021. https://arxiv.org/abs/2106.09672.
  41. 41.Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. Understanding dataset difficulty with v-usable information. In International Conference on Machine Learning (ICML), 2022. https://arxiv.org/abs/2110.08420.
  42. 42.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results, 2007. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html.
  43. 43.Sabri Eyuboglu, Bojan Karlaš, Christopher Ré, Ce Zhang, and James Zou. dcbench: a benchmark for data-centric ai systems. In Proceedings of the Sixth Workshop on Data Management for End-To-End Machine Learning, 2022. https://dl.acm.org/doi/abs/10.1145/3533028.3533310.
  44. 44.Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt. Data determines distributional robustness in contrastive language image pre-training (clip). In International Conference on Machine Learning (ICML), 2022. https://arxiv.org/abs/2205.01397.
  45. 45.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories. Conference on Computer Vision and Pattern Recognition (CVPR) Workshop, 2004. https://ieeexplore.ieee.org/document/1384978.
  46. 46.Dan Feldman, Matthew Faulkner, and Andreas Krause. Scalable training of mixture models via coresets. In Advances in Neural Information Processing Systems (NeuIPS), 2011. https://proceedings.neurips.cc/paper_files/paper/2011/file/2b6d65b9a9445c4271ab9076ead5605a-Paper.pdf.
  47. 47.Daniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper, Kayvon Fatahalian, and Christopher Ré. Fast and three-rious: Speeding up weak supervision with triplet methods. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/2002.11955.
  48. 48.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012. https://ieeexplore.ieee.org/abstract/document/6248074.
  49. 49.Amirata Ghorbani and James Zou. Data shapley: Equitable valuation of data for machine learning. In International Conference on Machine Learning, pp. 2242–2251. PMLR, 2019.
  50. 50.Stephan Graf and Olaf Mextorf. Just: Large-scale multi-tier storage infrastructure at the jülich supercomputing centre. Journal of large-scale research facilities JLSRF, 2021. https://jlsrf.org/index.php/lsf/article/view/180.
  51. 51.Chengcheng Guo, Bo Zhao, and Yanbing Bai. Deepcore: A comprehensive library for coreset selection in deep learning, 2022. https://arxiv.org/abs/2204.08499.
  52. 52.Han Guo, Nazneen Fatema Rajani, Peter Hase, Mohit Bansal, and Caiming Xiong. Fastif: Scalable influence functions for efficient model interpretation and debugging, 2020. https://arxiv.org/abs/2012.15781.
  53. 53.Jia Guo, Jiankang Deng, Alexandros Lattas, and Stefanos Zafeiriou. Sample and computation redistribution for efficient face detection. In International Conference on Learning Representations (ICLR), 2021. https://arxiv.org/abs/2105.04714.
  54. 54.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  55. 55.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. Annotation artifacts in natural language inference data. In Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2018. https://aclanthology.org/N18-2017.
  56. 56.Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi. Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023. https://arxiv.org/abs/2303.08114.
  57. 57.Frank R Hampel. The influence curve and its role in robust estimation. Journal of the american statistical association, 1974. https://www.jstor.org/stable/2285666.
  58. 58.Xiaochuang Han, Byron C Wallace, and Yulia Tsvetkov. Explaining black box predictions and unveiling data artifacts through influence functions, 2020. https://arxiv.org/abs/2005.06676.
  59. 59.A. Hanna, Emily L. Denton, Andrew Smart, and Jamila Smith-Loud. Towards a critical race methodology in algorithmic fairness. In Conference on Fairness, Accountability, and Transparency (FAccT), 2020. https://arxiv.org/abs/1912.03593.
  60. 60.Laura Hanu and Unitary team. Detoxify, 2020. https://github.com/unitaryai/detoxify.
  61. 61.Sariel Har-Peled and Soham Mazumdar. On coresets for k-means and k-median clustering. In Symposium on Theory of Computing (STOC), 2004. https://doi.org/10.1145/1007352.1007400.
  62. 62.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. https://arxiv.org/abs/1512.03385.
  63. 63.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2019. https://arxiv.org/abs/1709.00029.
  64. 64.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021. https://arxiv.org/abs/2006.16241.
  65. 65.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. https://arxiv.org/abs/1907.07174.
  66. 66.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models, 2022. https://arxiv.org/abs/2203.15556.
  67. 67.Raphael Hoffmann, Congle Zhang, Xiao Ling, Luke Zettlemoyer, and Daniel S Weld. Knowledge-based weak supervision for information extraction of overlapping relations. In Annual Meeting of the Association for Computational Linguistics (ACL), 2011. https://aclanthology.org/P11-1055.
  68. 68.Andrew Hundt, William Agnew, Vicky Zeng, Severin Kacianka, and Matthew Gombolay. Robots enact malignant stereotypes. In Conference on Fairness, Accountability, and Transparency (FAccT), 2022. https://arxiv.org/abs/2207.11569.
  69. 69.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. OpenCLIP, July 2021. https://doi.org/10.5281/zenodo.5143773.
  70. 70.Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt. Patching open-vocabulary models by interpolating weights. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://arXiv.org/abs/2208.05592.
  71. 71.Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. Datamodels: Predicting predictions from training data, 2022. https://arxiv.org/abs/2202.00622.
  72. 72.Tanuj Jain, Christopher Lennan, Zubin John, and Dat Tran. Imagededup, 2019. https://github.com/idealo/imagededup.
  73. 73.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2102.05918.
  74. 74.Mon-Fong Jiang, Shian-Shyong Tseng, and Chih-Ming Su. Two-phase clustering process for outliers detection. Pattern recognition letters, 2001. https://www.sciencedirect.com/science/article/abs/pii/S0167865500001318.
  75. 75.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 2019. https://arxiv.org/abs/1702.08734.
  76. 76.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. Conference on Computer Vision and Pattern Recognition (CVPR), 2017. https://arxiv.org/abs/1612.06890.
  77. 77.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification. In Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2017. https://arxiv.org/abs/1607.01759.
  78. 78.Juelich Supercomputing Center. JUWELS Booster Supercomputer, 2020. https://apps.fz-juelich.de/jsc/hps/juwels/configuration.html#hardware-configuration-of-the-system-name-booster-module.
  79. 79.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. https://arxiv.org/abs/2001.08361.
  80. 80.Kimmo Karkkainen and Jungseock Joo. Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. In IEEE/CVF Winter Conference on Applications of Computer Vision, 2021. https://arxiv.org/abs/1908.04913.
  81. 81.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. In International Conference on Machine Learning (ICML), 2017. https://arxiv.org/abs/1703.04730.
  82. 82.Pang Wei Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang. On the accuracy of influence functions for measuring group effects. Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1905.13289.
  83. 83.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. WILDS: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2012.07421.
  84. 84.Simon Kornblith, Jonathon Shlens, and Quoc V Le. Do better imagenet models transfer better? In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. https://arxiv.org/abs/1805.08974.
  85. 85.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In International Conference on Computer Vision Workshops (ICML), 2013. https://www.cv-foundation.org/openaccess/content_iccv_workshops_2013/W19/html/Krause_3D_Object_Representations_2013_ICCV_paper.html.
  86. 86.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images, 2009. https://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf.
  87. 87.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2012. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf.
  88. 88.Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. Adversarial filters of dataset biases. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/2002.04108.
  89. 89.Yann LeCun. The MNIST database of handwritten digits, 1998. http://yann.lecun.com/exdb/mnist/.
  90. 90.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. In Annual Meeting of the Association for Computational Linguistics (ACL), 2021. https://arxiv.org/abs/2107.06499.
  91. 91.Yi Li and Nuno Vasconcelos. Repair: Removing representation bias by dataset resampling. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. https://arxiv.org/abs/1904.07911.
  92. 92.Yulong Liu, Guibo Zhu, Bin Zhu, Qi Song, Guojing Ge, Haoran Chen, GuanHui Qiao, Ru Peng, Lingxiang Wu, and Jinqiao Wang. Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/6a386d703b50f1cf1f61ab02a15967bb-Paper-Datasets_and_Benchmarks.pdf.
  93. 93.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Conference on Computer Vision and Pattern Recognition (CVPR), 2022. https://arxiv.org/abs/2201.03545.
  94. 94.Mario Lucic, Matthew Faulkner, Andreas Krause, and Dan Feldman. Training gaussian mixture models at scale via coresets. Journal of Machine Learning Research (JMLR), 2018. http://jmlr.org/papers/v18/15-506.html.
  95. 95.S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi. Fine-grained visual classification of aircraft, 2013. https://arxiv.org/abs/1306.5151.
  96. 96.Gideon S Mann and Andrew McCallum. Generalized expectation criteria for semi-supervised learning with weakly labeled data. Journal of Machine Learning Research (JMLR), 2010. https://www.jmlr.org/papers/v11/mann10a.html.
  97. 97.Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Juan Ciro, Lora Aroyo, Bilge Acun, Sabri Eyuboglu, Amirata Ghorbani, Emmett Goodman, Tariq Kane, Christine R. Kirkpatrick, Tzu-Sheng Kuo, Jonas Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Zou, Carole-Jean Wu, Cody Coleman, Andrew Ng, Peter Mattson, and Vijay Janapa Reddi. Dataperf: Benchmarks for data-centric ai development, 2022. https://arxiv.org/abs/2207.10062.
  98. 98.Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/1906.01827.
  99. 99.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. In Advances in Neural Information Processing Systems (NeurIPS) Workshops, 2011. https://storage.googleapis.com/pub-tools-public-publication-data/pdf/37648.pdf.
  100. 100.Andrew Ng, Dillon Laird, and Lynn He. Data-centric ai competition, 2021. https://https-deeplearning-ai.github.io/data-centric-comp/.
  101. 101.Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, and Ludwig Schmidt. Quality not quantity: On the interaction between dataset design and robustness of clip. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://openreview.net/forum?id=LTCBavFWp5C.
  102. 102.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In Indian Conference on Computer Vision, Graphics and Image Processing, 2008. https://ieeexplore.ieee.org/document/4756141.
  103. 103.OpenAI. Gpt-4 technical report, 2023. https://arxiv.org/abs/2303.08774.
  104. 104.Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. Im2text: Describing images using 1 million captioned photographs. In Advances in Neural Information Processing Systems (NeurIPS), 2011. https://papers.nips.cc/paper_files/paper/2011/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-Paper.pdf.
  105. 105.Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. Cats and dogs. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012. https://ieeexplore.ieee.org/document/6248092.
  106. 106.Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. Deep learning on a data diet: Finding important examples early in training. In Advances in Neural Information Processing Systems (NeurIPS), 2021. https://arxiv.org/abs/2107.07075.
  107. 107.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V. Le. Combined scaling for zero-shot transfer learning, 2021. https://arxiv.org/abs/2111.10050.
  108. 108.Vinay Uday Prabhu and Abeba Birhane. Large image datasets: A pyrrhic win for computer vision? In Winter Conference on Applications of Computer Vision (WACV), 2020. https://arxiv.org/abs/2006.16923.
  109. 109.Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. Estimating training data influence by tracing gradient descent. Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2002.08484.
  110. 110.Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov, Simon Vandenhende, Yash Patel, Yi Wen, Vignesh Ramanathan, and Dhruv Mahajan. Filtering, distillation, and hard negatives for vision-language pre-training. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. https://arxiv.org/abs/2301.02280.
  111. 111.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2103.00020.
  112. 112.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision, 2022. https://arxiv.org/abs/2212.04356.
  113. 113.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research (JMLR), 2020. https://arxiv.org/abs/1910.10683.
  114. 114.Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron B. Adcock, Laurens van der Maaten, Deepti Ghadiyaram, and Olga Russakovsky. Beyond web-scraping: Crowd-sourcing a geodiverse datase, 2023. https://arxiv.org/abs/2301.02560.
  115. 115.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2102.12092.
  116. 116.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022. https://arxiv.org/abs/2204.06125.
  117. 117.A. J. Ratner, B. Hancock, J. Dunnmon, F. Sala, S. Pandey, and C. Ré. Training complex models with multi-task weak supervision. In Association for the Advancement of Artificial Intelligence (AAAI), 2019. https://arxiv.org/abs/1810.02840.
  118. 118.Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. Data programming: Creating large training sets, quickly. In Advances in Neural Information Processing Systems (NeurIPS), 2016. https://arxiv.org/abs/1605.07723.
  119. 119.Alexander J Ratner, Stephen H Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré. Snorkel: Rapid training data creation with weak supervision. In Very Large Data Bases Conference (VLDB), 2017. https://arxiv.org/abs/1711.10160.
  120. 120.Christopher Ré. Overton: A data system for monitoring and improving machine-learned products. In 10th Conference on Innovative Data Systems Research, CIDR 2020, Amsterdam, The Netherlands, January 12-15, 2020, Online Proceedings. www.cidrdb.org, 2020. URL http://cidrdb.org/cidr2020/papers/p33-re-cidr20.pdf.
  121. 121.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In International Conference on Machine Learning (ICML), 2019. http://proceedings.mlr.press/v97/recht19a.html.
  122. 122.William A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman. The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2022. https://openreview.net/forum?id=qnfYsave0U4.
  123. 123.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. https://arxiv.org/abs/2112.10752.
  124. 124.Peter J Rousseeuw and Mia Hubert. Robust statistics for outlier detection. Wiley interdisciplinary reviews: Data mining and knowledge discovery, 2011. http://i2pc.es/coss/Docencia/SignalProcessingReviews/Rousseeuw2011.pdf.
  125. 125.Peter J Rousseeuw and Mia Hubert. Anomaly detection by robust statistics. Wiley interdisciplinary reviews: Data mining and knowledge discovery, 2018. https://wires.onlinelibrary.wiley.com/doi/pdf/10.1002/widm.1236.
  126. 126.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 2015. https://arxiv.org/abs/1409.0575.
  127. 127.Shiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao, Sang Michael Xie, Kendrick Shen, Ananya Kumar, Weihua Hu, Michihiro Yasunaga, Henrik Marklund, Sara Beery, Etienne David, Ian Stavness, Wei Guo, Jure Leskovec, Kate Saenko, Tatsunori Hashimoto, Sergey Levine, Chelsea Finn, and Percy Liang. Extending the wilds benchmark for unsupervised adaptation. In International Conference on Learning Representations (ICLR), 2022. https://arxiv.org/abs/2112.05090.
  128. 128.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open dataset of clip-filtered 400 million image-text pairs, 2021. https://arxiv.org/abs/2111.02114.
  129. 129.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: An open large-scale dataset for training next generation image-text models. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2022. https://openreview.net/forum?id=M3Y74vmsMcY.
  130. 130.Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations (ICLR), 2018. https://openreview.net/forum?id=H1aIuk-RW.
  131. 131.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Annual Meeting of the Association for Computational Linguistics (ACL), 2018. https://aclanthology.org/P18-1238/.
  132. 132.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks?, 2021. https://arxiv.org/abs/2107.06383.
  133. 133.Changho Shin, Winfred Li, Harit Vishwakarma, Nicholas Roberts, and Frederic Sala. Universalizing weak supervision. In International Conference on Learning Representations (ICLR), 2022. https://openreview.net/forum?id=YpPiNigTzMT.
  134. 134.Haoyu Song, Li Dong, Weinan Zhang, Ting Liu, and Furu Wei. CLIP models are few-shot learners: Empirical studies on VQA and visual entailment. In Annual Meeting of the Association for Computational Linguistics (ACL), 2022. https://aclanthology.org/2022.acl-long.421.
  135. 135.Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos. Beyond neural scaling laws: beating power law scaling via data pruning. In Advances in Neural Information Processing Systems (NeurIPS), 2022. https://openreview.net/forum?id=UmvSlP-PyV.
  136. 136.Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021. https://arxiv.org/abs/2103.01913.
  137. 137.Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel. The german traffic sign recognition benchmark: a multi-class classification competition. In International Joint Conference on Neural Networks (IJCNN), 2011. https://ieeexplore.ieee.org/document/6033395.
  138. 138.Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020. https://aclanthology.org/2020.emnlp-main.746.
  139. 139.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://dl.acm.org/doi/abs/10.5555/3495724.3497285.
  140. 140.Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. YFCC100M: The new data in multimedia research. Communications of the ACM, 2016. https://arxiv.org/abs/1503.01817.
  141. 141.Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. An empirical study of example forgetting during deep neural network learning. In International Conference on Learning Representations (ICLR), 2018. https://arxiv.org/abs/1812.05159.
  142. 142.Bastiaan S Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling. Rotation equivariant CNNs for digital pathology, 2018. https://arxiv.org/abs/1806.03962.
  143. 143.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1905.13549.
  144. 144.Ryan Webster, Julien Rabin, Loic Simon, and Frederic Jurie. On the de-duplication of laion-2b, 2023. https://arxiv.org/abs/2303.12733.
  145. 145.Kai Wei, Rishabh Iyer, and Jeff Bilmes. Submodularity in data subset selection and active learning. In International Conference on Machine Learning (ICML), 2015. https://proceedings.mlr.press/v37/wei15.html.
  146. 146.Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva. Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision (IJCV), 2016. https://link.springer.com/article/10.1007/s11263-014-0748-y.
  147. 147.Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky. Towards fairer datasets: filtering and balancing the distribution of the people subtree in the imagenet hierarchy. In Conference on Fairness, Accountability, and Transparency (FAccT), 2020. https://arxiv.org/abs/1912.07726.
  148. 148.Kaiyu Yang, Jacqueline H Yau, Li Fei-Fei, Jia Deng, and Olga Russakovsky. A study of face obfuscation in ImageNet. In International Conference on Machine Learning (ICML), 2022. https://arxiv.org/abs/2103.06191.
  149. 149.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. Filip: Fine-grained interactive language-image pre-training. In International Conference on Learning Representations (ICLR), 2022. https://arxiv.org/abs/2111.07783.
  150. 150.Shuhei Yokoo. Contrastive learning with large memory bank and negative embedding subtraction for accurate copy detection, 2021. https://arxiv.org/abs/2112.04323.
  151. 151.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014. https://aclanthology.org/Q14-1006/.
  152. 152.Dantong Yu, Gholamhosein Sheikholeslami, and Aidong Zhang. Findout: Finding outliers in very large datasets. Knowledge and information Systems, 2002. https://link.springer.com/article/10.1007/s101150200013.
  153. 153.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision, 2021. https://arxiv.org/abs/2111.11432.
  154. 154.Man-Ching Yuen, Irwin King, and Kwong-Sak Leung. A survey of crowdsourcing systems. In SocialCom. IEEE, 2011. https://ieeexplore.ieee.org/document/6113213.
  155. 155.Matei Zaharia, Reynold S Xin, Patrick Wendell, Tathagata Das, Michael Armbrust, Ankur Dave, Xiangrui Meng, Josh Rosen, Shivaram Venkataraman, Michael J Franklin, et al. Apache spark: a unified engine for big data processing. Communications of the ACM, 2016. https://dl.acm.org/doi/10.1145/2934664.
  156. 156.Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, André Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby. The visual task adaptation benchmark, 2019. http://arxiv.org/abs/1910.04867.
  157. 157.Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang, and Alexander Ratner. WRENCH: A comprehensive benchmark for weak supervision. In NeurIPS, 2021. URL https://openreview.net/forum?id=Q9SKS5k8io.
  158. 158.Jieyu Zhang, Cheng-Yu Hsieh, Yue Yu, Chao Zhang, and Alexander Ratner. A survey on programmatic weak supervision, 2022. https://arxiv.org/abs/2202.05433.
  159. 159.Zhifei Zhang, Yang Song, and Hairong Qi. Age progression/regression by conditional adversarial autoencoder. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017. https://arxiv.org/abs/1702.08423.
  160. 160.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In European Conference on Computer Vision (ECCV), 2022. https://arxiv.org/abs/2201.02605.

Citation

MLA
Gadre, S. Y., et al. “DataComp: In Search of the Next Generation of Multimodal Datasets”. arXiv, 2023, http://arxiv.org/abs/2304.14108v5.
APA
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., Orgad, E., Entezari, R., Daras, G., Pratt, S., Ramanujan, V., Bitton, Y., Marathe, K., Mussmann, S., Vencu, R., … Schmidt, L. (2023). DataComp: In search of the next generation of multimodal datasets. arXiv. http://arxiv.org/abs/2304.14108v5
Chicago
Gadre, S. Y., G. Ilharco, A. Fang, et al. 2023. “DataComp: In Search of the Next Generation of Multimodal Datasets”. arXiv. http://arxiv.org/abs/2304.14108v5.
Harvard
Gadre, S.Y. et al. (2023) “DataComp: In search of the next generation of multimodal datasets”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.14108v5.
Vancouver
1. Gadre SY, Ilharco G, Fang A, et al (2023) DataComp: In search of the next generation of multimodal datasets. arXiv

BibTeX

@article{gadre2023datacomp,
  title = {DataComp: In search of the next generation of multimodal datasets},
  author = {Gadre, Samir Yitzhak and Ilharco, Gabriel and Fang, Alex and Hayase, Jonathan and Smyrnis, Georgios and Nguyen, Thao and Marten, Ryan and Wortsman, Mitchell and Ghosh, Dhruba and Zhang, Jieyu and Orgad, Eyal and Entezari, Rahim and Daras, Giannis and Pratt, Sarah and Ramanujan, Vivek and Bitton, Yonatan and Marathe, Kalyani and Mussmann, Stephen and Vencu, Richard and Cherti, Mehdi and Krishna, Ranjay and Koh, Pang Wei and Saukh, Olga and Ratner, Alexander and Song, Shuran and Hajishirzi, Hannaneh and Farhadi, Ali and Beaumont, Romain and Oh, Sewoong and Dimakis, Alex and Jitsev, Jenia and Carmon, Yair and Shankar, Vaishaal and Schmidt, Ludwig},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.14108v5},
  eprint = {2304.14108}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission