Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

Alex FangGabriel IlharcoMitchell WortsmanYuhao WanVaishaal ShankarAchal DaveLudwig Schmidt

article2022ICML195 citations

Demonstrates through systematic experiments that CLIP's resilience to natural distribution shifts stems almost entirely from training data diversity rather than language supervision, contrastive loss, or model scaling.

Listen

Machine learning vision models often suffer severe performance drops when deployed in real-world settings that differ from their training data. Recently, contrastive language-image pre-training (CLIP) models achieved unprecedented resilience against natural distribution shifts, such as changes in image style, background, or viewpoints. However, because CLIP introduced multiple technical changes simultaneously—including multimodal text-image supervision, massive web-scale data, specialized loss functions, and prompt-based testing—the primary cause of this improved robustness remained unresolved.

The article systematically evaluates five potential causes for CLIP's robustness: dataset size, training data distribution, language supervision during training, prompt design at testing time, and contrastive training objectives. Its core objective is to identify which specific component drives reliable performance under real-world data shifts.

To isolate these factors, the authors conducted controlled experiments across standard benchmark distributions (such as ImageNet) and challenging out-of-distribution test sets (including ImageNetV2, ImageNet-R, ImageNet-Sketch, and ObjectNet). They created ImageNet-Captions, a new dataset pairing over 463,000 standard benchmark images with their original natural language Flickr metadata, enabling a direct head-to-head comparison between language-image training and traditional single-label classification on identical images. Additionally, they introduced a simplified baseline on the 15-million-image YFCC dataset that used standard self-supervised visual pre-training paired with simple keyword matching, entirely removing complex language model architectures.

The investigation produced three central findings. First, training data distribution is the primary driver of out-of-distribution robustness. Models trained on diverse web data consistently demonstrated superior robustness compared to models trained on standard curated benchmarks. Second, natural language supervision during training does not inherently boost robustness; when trained on identical image sets, language-guided models performed no better against distribution shifts than standard image classifiers. Third, neither test-time prompt variations nor contrastive loss functions independently produced meaningful robustness gains, and prior work already established that training dataset size alone does not change effective robustness.

These results demonstrate that language supervision serves primarily as an efficient mechanism for aggregating broad, diverse web data without requiring manual labeling, rather than acting as an algorithmic cure for model fragility. For organizational decision-makers, this shifts strategic focus: investing in sophisticated language-vision architectures or prompt engineering will not solve reliability issues if the underlying image training distribution lacks real-world diversity. Robustness in computer vision is fundamentally a data-centric challenge rather than a modeling artifact.

Organizations developing or deploying reliable vision systems should prioritize data collection diversity and dataset curation over complex loss formulations or extensive prompt tuning. Future research and development should focus on identifying which specific visual properties within web distributions confer robustness and developing efficient methods to curate diverse visual data. While these conclusions are strongly supported across multiple standard benchmarks and model architectures, the findings are bounded by the specific natural distribution shifts tested, and further work is needed to determine whether specialized language-informed training can improve performance on narrower, domain-specific vision tasks.

arXiv: 2205.01397
Cover for Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

Abstract

Contrastively trained language-image models such as CLIP, ALIGN, and BASIC have demonstrated unprecedented robustness to multiple challenging natural distribution shifts. Since these language-image models differ from previous training approaches in several ways, an important question is what causes the large robustness gains. We answer this question via a systematic experimental investigation. Concretely, we study five different possible causes for the robustness gains: (i) the training set size, (ii) the training distribution, (iii) language supervision at training time, (iv) language supervision at test time, and (v) the contrastive loss function. Our experiments show that the more diverse training distribution is the main cause for the robustness gains, with the other factors contributing little to no robustness. Beyond our experimental results, we also introduce ImageNet-Captions, a version of ImageNet with original text annotations from Flickr, to enable further controlled experiments of language-image training.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Experimental setup for measuring robustness
  • 2.2. Additional related work
  • 3. ImageNet-Captions
  • 3.1. Constructing ImageNet-Captions
  • 3.2. Properties of ImageNet-Captions
  • 4. ImageNet-Captions experiments
  • 4.1. Caption construction
  • 4.2. Robustness
  • 4.3. Pre-training on language
  • 4.4. Effect of using templates
  • 4.5. Improving ImageNet performance using captions
  • 5. YFCC experiments
  • Effect of pre-trained language encoder
  • 6. Effect of test time prompts
  • 7. Effect of contrastive loss functions
  • 8. Conclusion
  • 9. Acknowledgements
  • References
  • A. Distribution shift examples
  • B. ImageNet-Captions experiments training details
  • C. ImageNet-Captions CLIP subsampled experiments
  • D. ImageNet-Captions classification experiments
  • E. ImageNet-Captions language encoder experiments
  • F. ImageNet-Captions template experiments
  • G. ImageNet templates and captions training variation experiments
  • H. Self-supervised training variation experiments
  • I. Data cleaning
  • J. NoCLIP ablations
  • K. YFCC-15M-Cls additional experiments
  • L. YFCC-15M-Cls classes
  • M. ImageNet-Captions additional statistics

Knowls

  1. Knowl 1 — Training Distribution as the Primary Driver of Effective Robustness in Contrastive Language-Image Models

    empirical result

    An empirical investigation into the source of the unprecedented robustness of contrastive language-image pre-trained models (such as CLIP) shows that robustness gains are governed almost exclusively by the choice of training data distribution. Specifically:

    1. Language supervision during training does not increase effective robustness: models trained on identical image sets with text captions versus discrete class labels follow the same linear robustness trend.
    2. Contrastive loss functions (e.g., SimCLRv2, SimSiam, SwAV) pre-trained on ImageNet do not improve effective robustness over standard supervised classifiers.
    3. Test-time natural language prompting variations shift accuracy but follow the same trajectory as interpolating with a random classifier, showing no intrinsic robustness benefit.
    4. Training set size alone does not change effective robustness, as sub-sampling studies confirm that size alters base accuracy but not out-of-distribution robustness trends.

    In contrast, training models on diverse web image distributions (such as YFCC-15M or OpenAI's web dataset) yields substantial improvements in effective robustness across natural distribution shifts (ImageNetV2, ImageNet-R, ImageNet-Sketch, and ObjectNet), regardless of whether training uses language-image contrastive learning or standard classification on pseudo-labeled classes.

  2. Knowl 2 — ImageNet-Captions Dataset

    experimental setup

    ImageNet-Captions is a multimodal dataset built to enable controlled comparisons between standard multiclass image classification and natural language-image supervision on an identical image distribution.

    Dataset construction:

    1. URL Matching: Starting from 14,197,122 image URLs in the ImageNet Fall 2011 release, images sourced from Flickr and corresponding to the 1,000 ILSVRC-2012 classes were extracted, yielding 642,147 images across 999 classes (excluding only the "teddy bear" synset).
    2. Metadata Retrieval: Original user metadata was queried via the Flickr API to collect title, description, and user-provided tags for each photo identifier.
    3. Deduplication and Cleaning: Deduplication against the ILSVRC-2012 training set and removal of captions containing profanity produced a final set of 463,622 image-text pairs with corresponding ImageNet class labels.

    Key characteristics:

    • Captions encompass 127 languages, with 89.9% in English; median caption length is 17 words.
    • 93.8% of the image captions explicitly contain the target ImageNet class label within the combined title, description, or tags (51.6% in title only, 28.9% in description only, and 73.8% in tags only).
  3. Knowl 3 — Equivalence of Effective Robustness Between Language-Image and Classification Models on ImageNet-Captions

    empirical result

    When evaluated across natural distribution shifts (ImageNetV2, ImageNet-R, ImageNet-Sketch, ObjectNet, and ImageNet-A), ResNet-50 models trained on ImageNet-Captions via CLIP contrastive learning follow the exact same linear effective robustness trend as standard classification models trained on the same image subset. Neither model exhibits the effective robustness lift observed in web-trained CLIP models.

    Experiment IN (%) IN-V2 (%) IN-R (%) IN Sketch (%) ObjectNet (%) IN-A (%)
    IN-Captions CLIP 30% 10.8 8.6 4.7 0.8 3.8 1.8
    IN-Captions CLIP 50% 18.9 14.2 7.4 1.2 5.7 2.2
    IN-Captions CLIP 100% 31.5 24.0 10.9 2.7 9.1 3.0
    IN-Captions Cls 30% 29.3 23.1 10.8 3.6 6.0 2.5
    IN-Captions Cls 60% 41.2 33.0 16.7 7.2 11.2 2.9
    IN-Captions Cls 100% 48.7 40.0 21.6 10.8 15.8 3.8
    IN-Captions Cls 100%, Aug 54.3 45.0 20.8 10.7 18.7 3.5

    All entries indicate top-1 accuracy (%). Because the visual distribution is controlled, these results demonstrate that multimodal language supervision at training time does not intrinsically confer out-of-distribution robustness.

  4. Knowl 4 — NoCLIP Baseline Training Procedure

    algorithm

    NoCLIP is a vision training framework that achieves the effective robustness of CLIP models without training or evaluating any language model.

    Input: Web image-caption dataset Dweb={(xi,ti)}i=1N\mathcal{D}_{\text{web}} = \{(x_i, t_i)\}_{i=1}^N, target class lexicon S\mathcal{S} containing class synsets and synonyms
    Output: Image classifier fθf_\theta
    1. Self-supervised visual pre-training:
       Train visual encoder backbone ϕ\phi on image inputs {xi}i=1N\{x_i\}_{i=1}^N using SimCLR, omitting all text captions tit_i.
    2. Classification dataset filtering:
       Initialize Dcls←∅\mathcal{D}_{\text{cls}} \leftarrow \emptyset
       for each (xi,ti)∈Dweb(x_i, t_i) \in \mathcal{D}_{\text{web}}:
           Find substring matches between text tit_i and lexicon entries in S\mathcal{S}.
           if text tit_i matches exactly one unique class synset c∈Sc \in \mathcal{S}:
               Add (xi,c)(x_i, c) to Dcls\mathcal{D}_{\text{cls}}
           else:
               Discard (xi,ti)(x_i, t_i)
    3. Supervised classification fine-tuning:
       Initialize a linear classification layer WW over ϕ\phi.
       Fine-tune ϕ\phi and WW on Dcls\mathcal{D}_{\text{cls}} using cross-entropy loss with class-balanced sampling, RandAugment data augmentation, cosine learning rate decay, and early stopping.

    When applied to YFCC-15M (N≈14.8MN \approx 14.8\text{M}), step 2 extracts 1,694,1251,694,125 images across 953 classes (11.4% of YFCC-15M), and fine-tuning runs for 1 epoch using a ViT-B/16 backbone.

  5. Knowl 5 — Performance and Robustness of NoCLIP vs. CLIP on YFCC-15M

    empirical result

    A ViT-B/16 model trained using the NoCLIP pipeline (SimCLR pre-training on YFCC-15M followed by 1 epoch of classification fine-tuning on the 1.69M filtered subset YFCC-15M-Cls) matches the accuracy and effective robustness of a CLIP model trained on all 15M image-caption pairs.

    Training Style ImageNet Top-1 (%) Average OOD Top-1 (%)
    CLIP 37.9 19.9
    NoCLIP (SimCLR →\to Classification) 35.7 18.8

    Average OOD accuracy is measured over four distribution shifts: ImageNetV2, ImageNet-R, ImageNet-Sketch, and ObjectNet. NoCLIP achieves an effective robustness trajectory nearly identical to CLIP while discarding 88.6% of the captions and completely avoiding language model training.

  6. Knowl 6 — Inefficacy of Self-Supervised Contrastive Losses Alone for Effective Robustness

    empirical result

    Pre-training vision models on ImageNet using self-supervised contrastive objectives—including SimCLRv2, SimSiam, and SwAV—does not yield effective robustness beyond standard supervised ImageNet models.

    When plotting in-distribution ImageNet top-1 accuracy against average accuracy over five distribution shifts (ImageNetV2, ImageNet-R, ImageNet-Sketch, ObjectNet, and ImageNet-A), contrastive models align closely with the standard ImageNet classification baseline fit. While non-contrastive Masked Autoencoders (MAE) exhibit a minor robustness improvement, they remain substantially below the zero-shot CLIP robustness line. Thus, contrastive loss formulations in isolation do not account for CLIP's distributional robustness.

  7. Knowl 7 — Effect of Test-Time Prompts on CLIP Robustness

    empirical result

    Systematic evaluation of over 100 prompt variations for CLIP—varying template text (standard CLIP templates, raw class names, random word affixes), class name lexicons (Radford et al. labels, WordNet synsets, combined), number of ensembled templates (11 to 8080), and number of synonyms per class (11 to 44)—shows that prompt design is not the origin of CLIP's effective robustness.

    Although specific prompt configurations alter absolute in-distribution and out-of-distribution top-1 accuracy, any shift in effective robustness matches the exact curve produced by linearly interpolating the model's predictions with a random classifier (which naturally incurs zero accuracy drop under distribution shift).

  8. Knowl 8 — Pre-trained Language Encoders Do Not Induce Effective Robustness on Fixed Image Distributions

    empirical result

    To evaluate whether language representations pre-trained on large-scale web text impart visual robustness, models were trained on ImageNet-Captions using a randomly initialized ResNet-50 vision encoder paired with the pre-trained text encoder from OpenAI's CLIP (trained on 400M web pairs), under both fine-tuned and frozen text encoder configurations.

    While initializing or freezing the text encoder with pre-trained weights improves base ImageNet top-1 accuracy (from 31.5% for scratch initialization up to 38.3% for frozen text features), neither variant provides additional effective robustness above the standard ImageNet classification baseline across ImageNetV2, ImageNet-R, ImageNet-Sketch, ObjectNet, or ImageNet-A.

  9. Knowl 9 — Impact of Metadata Fields and Language Filtering in ImageNet-Captions

    data/table

    Evaluating different caption combinations and language filtering on ImageNet-Captions using a ResNet-50 CLIP model shows that incorporating comprehensive metadata improves performance, whereas English-only filtering reduces accuracy due to dataset shrinkage.

    Title Desc Tags Filter Size Rel. Size (%) IN IN-V2 IN-R IN Sketch ObjectNet
    ✓ ✓ 197K 42.6 15.7 12.2 6.6 1.1 5.5
    ✓ 459K 99.0 26.2 20.7 9.5 2.6 8.4
    ✓ ✓ ✓ 312K 67.4 21.9 16.5 8.0 1.7 6.1
    ✓ ✓ 461K 99.4 27.8 21.6 9.6 3.0 8.0
    ✓ ✓ ✓ ✓ 367K 79.3 26.5 20.3 8.9 2.3 7.9
    ✓ ✓ ✓ 464K 100.0 31.5 24.0 10.9 2.7 9.1

    All dataset shift columns report top-1 accuracy (%). Filtering for English drops the dataset size by up to 57.4%, degrading accuracy across all shifts. Using unfiltered Title, Description, and Tags achieves the highest ImageNet top-1 accuracy of 31.5%.

  10. Knowl 10 — Effective Robustness Metric Formulation

    definition

    Effective robustness quantifies out-of-distribution performance gains relative to a baseline model family with matching in-distribution accuracy.

    Let D1D_1 denote an in-distribution evaluation set (such as the ImageNet ILSVRC-2012 validation set) and D2D_2 an out-of-distribution shift test set (such as ImageNetV2, ImageNet-R, ImageNet-Sketch, or ObjectNet). Let accD1(f)\text{acc}_{D_1}(f) and accD2(f)\text{acc}_{D_2}(f) represent the top-1 accuracy of a model ff on distributions D1D_1 and D2D_2, respectively.

    A baseline mapping β:R→R\beta : \mathbb{R} \to \mathbb{R} is fit to the distribution pairs (accD1(f),accD2(f))\left(\text{acc}_{D_1}(f), \text{acc}_{D_2}(f)\right) across a reference family of standard baseline models (e.g., standard supervised ImageNet classifiers). For an evaluated model f′f', effective robustness ρ(f′)\rho(f') is defined as: ρ(f′)=accD2(f′)−β(accD1(f′))\rho(f') = \text{acc}_{D_2}(f') - \beta\left(\text{acc}_{D_1}(f')\right) A model demonstrates positive effective robustness when ρ(f′)>0\rho(f') > 0, indicating accuracy on D2D_2 beyond what is predicted by its in-distribution accuracy on D1D_1.

Coverage note — Deliberately omitted the specific list of 47 unrepresented ImageNet synsets in YFCC-15M and the specific profanity keyword list used during data cleaning, as these are minor low-level implementation details.

References

  1. 1.Andreassen, A., Bahri, Y., Neyshabur, B., and Roelofs, R. The evolution of out-of-distribution robustness throughout fine-tuning. 2021. https://arxiv.org/abs/2106.15831.
  2. 2.Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://proceedings.neurips.cc/paper/2019/file/97af07a14cacba681feacf3012730892-Paper.pdf.
  3. 3.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. 2020. https://arxiv.org/abs/2006.09882.
  4. 4.Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. https://arxiv.org/abs/2102.08981.
  5. 5.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research. PMLR, 2020a. http://proceedings.mlr.press/v119/chen20j.html.
  6. 6.Chen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G. Big self-supervised models are strong semisupervised learners. 2020b. https://arxiv.org/abs/2006.10029.
  7. 7.Chen, X. and He, K. Exploring simple siamese representation learning. corr abs/2011.10566 (2020). 2020. https://arxiv.org/abs/2011.10566.
  8. 8.Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L. Microsoft coco captions: Data collection and evaluation server. 2015. https://arxiv.org/abs/1504.00325.
  9. 9.Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V. Randaugment: Practical automated data augmentation with a reduced search space. In Conference on Computer Vision and Pattern Recognition Workshops, 2020. https://arxiv.org/abs/1909.13719.
  10. 10.Desai, K. and Johnson, J. Virtex: Learning visual representations from textual annotations. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pp. 11162–11173. Computer Vision Foundation / IEEE, 2021. URL https://openaccess.thecvf.com/content/CVPR2021/html/Desai_VirTex_Learning_Visual_Representations_From_Textual_Annotations_CVPR_2021_paper.html.
  11. 11.Desai, K., Kaul, G., Aysola, Z. T., and Johnson, J. Redcaps: Web-curated image-text data created by the people, for the people. 2021. https://arxiv.org/abs/2111.11431.
  12. 12.Devillers, B., Choksi, B., Bielawski, R., and VanRullen, R. Does language help generalization in vision models? 2021. https://arxiv.org/abs/2104.08313.
  13. 13.Djolonga, J., Yung, J., Tschannen, M., Romijnders, R., Beyer, L., Kolesnikov, A., Puigcerver, J., Minderer, M., D’Amour, A., Moldovan, D., Gelly, S., Houlsby, N., Zhai, X., and Lucic, M. On robustness and transferability of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021. https://arxiv.org/abs/2007.08558.
  14. 14.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. CoRR, abs/2010.11929, 2020. https://arxiv.org/abs/2010.11929.
  15. 15.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 770–778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. https://doi.org/10.1109/CVPR.2016.90.
  16. 16.He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. Masked autoencoders are scalable vision learners. 2021. https://arxiv.org/abs/2111.06377.
  17. 17.Heckel, R. and Yilmaz, F. F. Early stopping in deep networks: Double descent and how to eliminate it. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=tlV90jvZbw.
  18. 18.Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D. Natural adversarial examples.(2019). 2019. https://arxiv.org/abs/1907.07174.
  19. 19.Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. https://arxiv.org/abs/2006.16241.
  20. 20.Jain, T., Lennan, C., John, Z., and Tran, D. Imagededup. https://github.com/idealo/imagededup, 2019.
  21. 21.Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision, 2021. https://arxiv.org/abs/2102.05918.
  22. 22.Miller, G. A. Wordnet: A lexical database for english. Commun. ACM, Nov 1995. https://doi.org/10.1145/219717.219748.
  23. 23.Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L. Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In International Conference on Machine Learning. PMLR, 2021. https://arxiv.org/abs/2107.04649.
  24. 24.Mu, N., Kirillov, A., Wagner, D., and Xie, S. Slip: Self-supervision meets language-image pre-training. 2021. https://arxiv.org/abs/2112.12750.
  25. 25.Ordonez, V., Kulkarni, G., and Berg, T. Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems, 24, 2011. https://proceedings.neurips.cc/paper/2011/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-Paper.pdf.
  26. 26.Pham, H., Dai, Z., Ghiasi, G., Liu, H., Yu, A. W., Luong, M., Tan, M., and Le, Q. V. Combined scaling for zero-shot transfer learning. CoRR, abs/2111.10050, 2021. https://arxiv.org/abs/2111.10050.
  27. 27.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research. PMLR, 2021. http://proceedings.mlr.press/v139/radford21a.html.
  28. 28.Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do imagenet classifiers generalize to imagenet? In International Conference on Machine Learning. PMLR, 2019. https://arxiv.org/abs/1902.10811.
  29. 29.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252, 2015.
  30. 30.Sariyildiz, M. B., Perez, J., and Larlus, D. Learning visual representations with caption annotations. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J. (eds.), Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VIII, volume 12353 of Lecture Notes in Computer Science, pp. 153–170. Springer, 2020. doi: 10.1007/978-3-030-58598-3_10. URL https://doi.org/10.1007/978-3-030-58598-3_10.
  31. 31.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. 2021. https://arxiv.org/abs/2111.02114.
  32. 32.Shankar, V., Roelofs, R., Mania, H., Fang, A., Recht, B., and Schmidt, L. Evaluating machine accuracy on imagenet. In International Conference on Machine Learning. PMLR, 2020. https://proceedings.mlr.press/v119/shankar20c.html.
  33. 33.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018.
  34. 34.Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. 2021. https://arxiv.org/abs/2103.01913.
  35. 35.Taori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L. Measuring robustness to natural distribution shifts in image classification. 2020. https://arxiv.org/abs/2007.00644.
  36. 36.Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016.
  37. 37.Wang, H., Ge, S., Xing, E. P., and Lipton, Z. C. Learning robust global representations by penalizing local predictive power. 2019. https://arxiv.org/abs/1905.13549.
  38. 38.Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. Lit: Zero-shot transfer with locked-image text tuning. CoRR, 2021. https://arxiv.org/abs/2111.07991.
  39. 39.Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., and Langlotz, C. P. Contrastive learning of medical visual representations from paired images and text. CoRR, abs/2010.00747, 2020. URL https://arxiv.org/abs/2010.00747.

Citation

MLA
Fang, A., et al. “Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”. International Conference on Machine Learning, vol. 162, 2022, pp. 6216–34, https://proceedings.mlr.press/v162/fang22a.html.
APA
Fang, A., Ilharco, G., Wortsman, M., Wan, Y., Shankar, V., Dave, A., & Schmidt, L. (2022). Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP). International Conference on Machine Learning, 162, 6216–6234. https://proceedings.mlr.press/v162/fang22a.html
Chicago
Fang, A., G. Ilharco, M. Wortsman, et al. 2022. “Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”. International Conference on Machine Learning 162: 6216–34. https://proceedings.mlr.press/v162/fang22a.html.
Harvard
Fang, A. et al. (2022) “Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”, International Conference on Machine Learning. PMLR, pp. 6216–6234. Available at: https://proceedings.mlr.press/v162/fang22a.html.
Vancouver
1. Fang A, Ilharco G, Wortsman M, Wan Y, Shankar V, Dave A, Schmidt L (2022) Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP). In: International Conference on Machine Learning. PMLR, pp 6216–6234

BibTeX

@InProceedings{pmlr-v162-fang22a,
  title = 	 {Data Determines Distributional Robustness in Contrastive Language Image Pre-training ({CLIP})},
  author =       {Fang, Alex and Ilharco, Gabriel and Wortsman, Mitchell and Wan, Yuhao and Shankar, Vaishaal and Dave, Achal and Schmidt, Ludwig},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {6216--6234},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/fang22a/fang22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/fang22a.html},
  abstract = 	 {Contrastively trained language-image models such as CLIP, ALIGN, and BASIC have demonstrated unprecedented robustness to multiple challenging natural distribution shifts. Since these language-image models differ from previous training approaches in several ways, an important question is what causes the large robustness gains. We answer this question via a systematic experimental investigation. Concretely, we study five different possible causes for the robustness gains: (i) the training set size, (ii) the training distribution, (iii) language supervision at training time, (iv) language supervision at test time, and (v) the contrastive loss function. Our experiments show that the more diverse training distribution is the main cause for the robustness gains, with the other factors contributing little to no robustness. Beyond our experimental results, we also introduce ImageNet-Captions, a version of ImageNet with original text annotations from Flickr, to enable further controlled experiments of language-image training.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/