Robust Cross-Modal Representation Learning with Progressive Self-Distillation

Alex AndonianShixing ChenRaffay Hamid

article2022CVPR71 citations

Proposes a progressive self-distillation framework that replaces rigid one-to-one pairings in vision-language pretraining with dynamic soft-alignment targets, consistently outperforming CLIP across zero-shot, transfer, and retrieval benchmarks without extra computational overhead.

Listen

Modern vision-language artificial intelligence models, such as Contrastive Language-Image Pretraining (CLIP), achieve impressive capabilities by training on hundreds of millions of web-scraped image-caption pairs. However, web data is inherently noisy and loosely descriptive. Traditional training frameworks strictly enforce one-to-one pairings, wrongly penalizing images that match multiple captions or captions that relate to multiple images. This rigid assumption causes high data and compute inefficiency, requiring thousands of processor days to learn effectively.

The article demonstrates a novel training framework that uses progressive self-distillation and soft alignment targets to learn robust multimodal representations directly from noisy data. The primary objective is to evaluate whether allowing flexible, many-to-many probability alignments improves model accuracy, robustness, and efficiency compared to standard contrastive methods without adding computational overhead.

The researchers evaluated their approach across three pretraining datasets of varying scale and noise, ranging from roughly 118,000 to 10 million image-text pairs. Rather than pruning or hand-filtering noisy data, the model acts as its own teacher. During each training batch, it dynamically splits samples into an aligned subset trained with standard objectives and an unaligned subset where the model predicts softened alignment targets. Over the course of training, the model progressively increases its reliance on these self-generated soft targets while utilizing cross-modal swapped predictions to avoid reinforcing its own errors. The resulting models were benchmarked across 14 standard test datasets across zero-shot classification, linear probe transfer, and cross-modal retrieval tasks.

The findings demonstrate substantial, consistent performance advantages over baseline CLIP models. Across zero-shot classification benchmarks, the proposed framework achieved an absolute average accuracy gain of 2.22% on smaller data, 6.19% on intermediate data, and 5.23% on large-scale data. On out-of-distribution robustness tests designed to measure how well models handle real-world variations and distribution shifts, the method surpassed baseline accuracy by up to 8.5%. Linear probe and image-text retrieval benchmarks similarly showed consistent improvements across all test domains. In addition, data efficiency sweeps revealed that the relative performance advantage over CLIP persisted and widened across data sizes spanning two orders of magnitude.

These results indicate that accounting for natural semantic overlap and noisy annotations significantly improves representation learning without requiring larger infrastructure budgets. For organizations deploying computer vision and multimodal search systems, this approach reduces computational costs, improves data utilization, and enhances operational reliability against real-world visual shifts. Because it avoids maintaining complex secondary teacher networks or separate data-filtering pipelines, the method can be integrated directly into existing training workflows.

Organizations training multimodal models should adopt dynamic soft alignments and progressive self-distillation in place of rigid contrastive objectives. Future development should explore applying these self-distillation techniques to larger foundation models, specialized architectures, and redundant dataset optimizations to further minimize hardware resource requirements.

While confidence in the comparative empirical improvements across the 14 benchmarks is high, the largest dataset evaluated in the article contained approximately 10 million pairs, which is smaller than private industry datasets containing hundreds of millions of samples. Stakeholders should conduct pilot evaluations when scaling up to multi-billion-parameter models or domain-specific enterprise datasets to verify that hyperparameter decay schedules transfer seamlessly.

arXiv: 2204.04588
Cover for Robust Cross-Modal Representation Learning with Progressive Self-Distillation

Abstract

The learning objective of vision-language approach of CLIP [63] does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets, which contributes to its compute and data inefficiency. To address this challenge, we introduce a novel training framework based on cross-modal contrastive learning that uses progressive self-distillation and soft image-text alignments to more efficiently learn robust representations from noisy data. Our model distills its own knowledge to dynamically generate soft-alignment targets for a subset of images and captions in every minibatch, which are then used to update its parameters. Extensive evaluation across 14 benchmark datasets shows that our method consistently outperforms its CLIP counterpart in multiple settings, including: (a) zero-shot classification, (b) linear probe transfer, and (c) image-text retrieval, without incurring extra computational cost. Analysis using an ImageNet-based robustness test-bed [70] reveals that our method offers better effective robustness to natural distribution shifts compared to both ImageNet-trained models and CLIP itself. Lastly, pretraining with datasets spanning two orders of magnitude in size shows that our improvements over CLIP tend to scale with number of training examples.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. Preliminaries
  • 3.2. Contrastive Learning with InfoNCE Loss
  • 3.3. Distillation through Soft-Alignments
  • 3.4. Progressive Self-Distillation
  • 3.4.1 Teacher Network Selection
  • 3.4.2 Progressing from Student to Teacher
  • 4. Experiments
  • 4.1. Pretraining Datasets
  • 4.2. Pretraining Details
  • 4.3. Evaluation Details
  • 4.4. Zero-Shot Image Classification
  • 4.5. Evaluation of Effective Robustness
  • 4.6. Linear Probe Performance
  • 4.7. Image-Text Retrieval
  • 4.8. Ablation Study
  • 4.9. Qualitative Analysis
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Progressive Self-Distillation Contrastive Learning Objective

    equation

    Let a batch of NN image-text pairs be represented by ℓ2\ell_2-normalized feature embedding matrices V,T∈RN×d\mathbf{V}, \mathbf{T} \in \mathbb{R}^{N \times d} produced by an image encoder fvf_v and a text encoder ftf_t. In Progressive Self-Distillation (PSD), the current state of the model acts as its own teacher (f~v=fv\tilde{f}_v = f_v, f~t=ft\tilde{f}_t = f_t), yielding teacher embedding matrices V~=V\tilde{\mathbf{V}} = \mathbf{V} and T~=T\tilde{\mathbf{T}} = \mathbf{T}.

    A batch is dynamically partitioned into Na=⌊αN⌋N_a = \lfloor \alpha N \rfloor aligned instances and Nu=N−⌊αN⌋N_u = N - \lfloor \alpha N \rfloor unaligned instances, where α∈[0,1]\alpha \in [0, 1] is a partitioning factor. Let V~a,T~a∈RNa×d\tilde{\mathbf{V}}_a, \tilde{\mathbf{T}}_a \in \mathbb{R}^{N_a \times d} denote the aligned teacher embeddings (the first NaN_a rows) and Vu,Tu∈RNu×d\mathbf{V}_u, \mathbf{T}_u \in \mathbb{R}^{N_u \times d} denote the unaligned student embeddings (the remaining NuN_u rows).

    The total pretraining loss LInfoNCEPSD\mathcal{L}_{\text{InfoNCE}}^{\text{PSD}} combines standard InfoNCE loss on aligned pairs with soft-distillation loss on unaligned pairs:

    LInfoNCEPSD=α[H(INa,ρ(V~aT~⊤;τ))+H(INa,ρ(T~aV~⊤;τ))]+(1−α)[H(Av,ρ(VuT⊤;τ))+H(At,ρ(TuV⊤;τ))]\mathcal{L}_{\text{InfoNCE}}^{\text{PSD}} = \alpha \left[ \mathcal{H}(\mathbf{I}_{N_a}, \rho(\tilde{\mathbf{V}}_a \tilde{\mathbf{T}}^\top; \tau)) + \mathcal{H}(\mathbf{I}_{N_a}, \rho(\tilde{\mathbf{T}}_a \tilde{\mathbf{V}}^\top; \tau)) \right] + (1 - \alpha) \left[ \mathcal{H}(\mathbf{A}_v, \rho(\mathbf{V}_u \mathbf{T}^\top; \tau)) + \mathcal{H}(\mathbf{A}_t, \rho(\mathbf{T}_u \mathbf{V}^\top; \tau)) \right]

    where:

    • H(P,Q)=−1M∑i=1M∑j=1NPijlog⁡Qij\mathcal{H}(\mathbf{P}, \mathbf{Q}) = -\frac{1}{M} \sum_{i=1}^M \sum_{j=1}^N P_{ij} \log Q_{ij} is the row-wise cross-entropy with mean reduction over MM rows.
    • ρ(Z;τ)ij=exp⁡(Zij/τ)∑k=1Nexp⁡(Zik/τ)\rho(\mathbf{Z}; \tau)_{ij} = \frac{\exp(Z_{ij}/\tau)}{\sum_{k=1}^N \exp(Z_{ik}/\tau)} is the row-wise softmax operator scaled by a learnable temperature τ\tau.
    • INa=[INa×Na∣0Na×Nu]∈RNa×N\mathbf{I}_{N_a} = [\mathbf{I}_{N_a \times N_a} \mid \mathbf{0}_{N_a \times N_u}] \in \mathbb{R}^{N_a \times N} is the zero-padded identity matrix.
    • Av∈RNu×N\mathbf{A}_v \in \mathbb{R}^{N_u \times N} and At∈RNu×N\mathbf{A}_t \in \mathbb{R}^{N_u \times N} are soft-alignment target probability distributions generated by the teacher corresponding to the unaligned student rows.
  2. Knowl 2 — Swapped Prediction for Cross-Modal Soft-Alignment Targets

    model/method

    To generate soft-alignment targets that supervise the student network without reinforcing model errors or causing representation collapse, the teacher network employs a swapped prediction strategy.

    Given the batch teacher embedding matrices V~,T~∈RN×d\tilde{\mathbf{V}}, \tilde{\mathbf{T}} \in \mathbb{R}^{N \times d} and a secondary teacher softmax temperature τ~\tilde{\tau}, the target alignment distributions across the batch are computed as:

    Av=ρ(T~V~⊤τ~)andAt=ρ(V~T~⊤τ~)\mathbf{A}_v = \rho\left(\frac{\tilde{\mathbf{T}}\tilde{\mathbf{V}}^\top}{\tilde{\tau}}\right) \quad \text{and} \quad \mathbf{A}_t = \rho\left(\frac{\tilde{\mathbf{V}}\tilde{\mathbf{T}}^\top}{\tilde{\tau}}\right)

    where ρ(⋅)\rho(\cdot) denotes the row-wise softmax operator.

    For the unaligned image student embeddings Vu\mathbf{V}_u, the image alignment target Av\mathbf{A}_v is computed from the text encoder's posterior probabilities over all images in the batch. Conversely, for the unaligned text student embeddings Tu\mathbf{T}_u, the text alignment target At\mathbf{A}_t is computed from the visual encoder's posterior probabilities over all captions in the batch. By aggregating information across all instances of the opposite modality, swapped prediction assigns soft, many-to-many probability weights Aij∈[0,1]A_{ij} \in [0, 1] to pairs (vi,tj)(v_i, t_j), re-calibrating false negatives and noisy pairings.

  3. Knowl 3 — Dynamic Minibatch Partitioning and Progressive Teacher Scheduling

    model/method

    Progressive self-distillation dynamically transitions the model from learning via hard ground-truth supervision to learning via its own soft-alignment predictions using two mechanisms:

    1. Dynamic Minibatch Partitioning: In each minibatch of NN image-text pairs, Na=⌊αN⌋N_a = \lfloor \alpha N \rfloor pairs are assigned as hard-grounded aligned instances and Nu=N−⌊αN⌋N_u = N - \lfloor \alpha N \rfloor pairs are assigned as soft-supervised unaligned instances. The partition is dynamic: assignments are randomly reshuffled at every training epoch rather than fixed globally, ensuring every sample receives both direct contrastive anchoring and soft-target refinement.
    2. Cosine Annealing Partitioning Schedule: The parameter α∈[0,1]\alpha \in [0, 1], which controls the proportion of hard vs. soft supervision in the loss function, is decayed across training iterations following a cosine annealing schedule from αstart=0.8\alpha_{\text{start}} = 0.8 down to αend=0.2\alpha_{\text{end}} = 0.2. Early in training when feature representations are noisy, the network relies on hard ground-truth labels (α=0.8\alpha = 0.8). As representations become more reliable, the teacher's soft alignment influence increases (1−α=0.81 - \alpha = 0.8), progressively turning the network into its own teacher.
  4. Knowl 4 — Pretraining Setup and Hyperparameters

    experimental setup

    The vision-language models are trained under the following configuration:

    • Vision Architecture: ViT-B/32 Vision Transformer receiving input images randomly cropped and resized to 224×224224 \times 224 pixels.
    • Text Architecture: Transformer text encoder with sequence lengths capped at 77 tokens via random sub-sequence sampling.
    • Embedding Dimension: Image and text representations are linearly projected to a shared 512-dimensional embedding space and normalized using the ℓ2\ell_2 norm.
    • Optimization: Adam optimizer with weight decay, trained from scratch for 100 epochs using automatic mixed-precision (AMP).
    • Batch Size and Hardware: Minibatch size of 4096 trained across up to 8 Nvidia A100 GPUs.
    • Temperature Parameter: Learnable InfoNCE temperature τ\tau is initialized to 0.07 and clamped to values ≤100\le 100.
    • Partitioning Schedule: Partition factor α\alpha is decayed from 0.8 to 0.2 using cosine annealing.
    • Pretraining Datasets:
      • MS COCO Captions: ∼\sim118K images, each with 5 human annotations.
      • Conceptual Captions 3M (CC3M): ∼\sim2.9M web-harvested image-alt-text pairs after preprocessing.
      • Conceptual Captions 12M (CC12M): ∼\sim10M image-alt-text pairs.
  5. Knowl 5 — Zero-Shot Image Classification Performance

    data/table

    Zero-shot Top-1 classification accuracy (%) of Progressive Self-Distillation (Ours) compared to baseline CLIP across 10 benchmark datasets, evaluating standard recognition (CIFAR-10, CIFAR-100, Caltech101, Places365, ImageNet) and robustness to natural distribution shifts (ObjectNet, ImageNet-R, ImageNet-O, ImageNet-A, ImageNetV2):

    Pretraining Dataset Method Cifar10 Cifar100 Caltech101 Places365 ObjectNet ImageNet-R ImageNet-O ImageNet-A ImageNetV2 ImageNet Average
    COCO CLIP 64.14 19.57 32.88 12.78 4.98 8.27 8.00 3.32 7.41 8.18 16.87
    COCO Ours 66.74 24.49 34.26 14.15 6.18 11.25 9.85 5.32 8.99 9.49 19.07 (+2.22)
    CC3M CLIP 73.90 30.60 54.07 24.54 4.49 28.33 18.50 7.827 21.43 23.56 28.73
    CC3M Ours 80.15 38.27 64.45 28.07 9.21 37.31 26.20 10.81 26.70 27.96 34.91 (+6.19)
    CC12M CLIP 75.29 41.94 75.86 31.29 12.50 51.34 31.90 13.25 34.89 37.87 40.61
    CC12M Ours 84.84 51.34 80.00 34.08 15.24 59.29 33.00 18.85 39.16 42.24 45.85 (+5.23)

    Progressive self-distillation outperforms baseline CLIP across all evaluated pretraining datasets and downstream classification tasks, yielding average zero-shot top-1 improvements of +2.22% on COCO, +6.19% on CC3M, and +5.23% on CC12M. Improvements are especially pronounced on out-of-distribution benchmarks such as ImageNet-R (+8.98% on CC3M and +7.95% on CC12M).

  6. Knowl 6 — Linear Probe Classification of Extracted Visual Features

    data/table

    To assess within-modal visual representation quality, a linear classifier is trained with L-BFGS on frozen visual features extracted by the image encoder across four downstream benchmark datasets. Top-1 linear probe accuracy (%):

    Pretraining Dataset Method Food101 OxfordPets Birdsnap ImageNet
    COCO CLIP 53.04 76.86 37.17 52.66
    COCO Ours 53.60 80.59 42.85 56.21
    CC3M CLIP 53.33 78.11 37.75 56.28
    CC3M Ours 60.69 80.40 43.17 61.00
    CC12M CLIP 67.41 85.17 41.06 59.42
    CC12M Ours 71.87 86.32 47.16 65.33

    Visual encoders trained with progressive self-distillation consistently outperform their CLIP counterparts across all four classification tasks and pretraining dataset sizes (e.g., ImageNet linear probe accuracy improves by +3.55% for COCO, +4.72% for CC3M, and +5.91% for CC12M), demonstrating that cross-modal soft distillation enhances uni-modal visual representations.

  7. Knowl 7 — Cross-Modal Image-Text Retrieval on MS COCO

    data/table

    Zero-shot and finetuned cross-modal retrieval performance evaluated on the MS COCO 5K test set across text-to-image and image-to-text subtasks, reported via Recall at KK (R@K↑R@K \uparrow in %) and Mean Rank (MnR↓\text{MnR} \downarrow). Asterisk (∗^*) indicates models pretrained on Conceptual Captions and subsequently finetuned on the MS COCO training set:

    Text-to-Image Image-to-Text
    Pretraining Dataset Method R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MnR ↓\downarrow R@1 ↑\uparrow R@5 ↑\uparrow R@10 ↑\uparrow MnR ↓\downarrow
    COCO CLIP 27.76 57.34 70.70 23.10 27.40 56.31 68.65 20.11
    COCO Ours 28.42 57.14 68.86 26.11 28.53 56.75 68.10 22.01
    CC3M CLIP 12.50 29.76 40.92 91.04 9.88 24.86 35.30 106.94
    CC3M Ours 16.98 37.12 48.28 63.80 13.19 31.54 43.00 72.92
    CC12M CLIP 19.64 40.66 51.72 55.23 17.63 39.67 50.77 55.29
    CC12M Ours 22.94 46.60 57.82 46.24 22.79 45.95 56.81 43.41
    CC3M CLIP∗^* 31.30 59.54 71.80 21.15 29.18 58.66 70.23 17.02
    CC3M Ours∗^* 33.26 66.70 76.92 18.77 32.20 60.14 71.96 17.05
    CC12M CLIP∗^* 35.22 62.74 73.46 17.70 33.36 61.15 73.36 15.76
    CC12M Ours∗^* 38.66 66.74 77.10 13.35 38.15 65.85 77.02 12.54

    Progressive self-distillation outperforms baseline CLIP on both retrieval directions across pretraining scales. In zero-shot retrieval on CC12M, R@1 increases by +3.30% (text-to-image) and +5.16% (image-to-text), with corresponding substantial decreases in mean rank.

  8. Knowl 8 — Ablation of Soft-Alignment Prediction Mode, Partitioning, and Distillation Scheduling

    data/table

    Ablation study evaluating the isolated contributions of target prediction mode (Forward Bootstrapping vs. Swapped Prediction), minibatch partitioning (Static vs. Dynamic), and self-distillation scheduling (Static α\alpha vs. Progressive Cosine Annealing) on models pretrained on MS COCO. Performance is evaluated via zero-shot ImageNet Top-1 accuracy (%) and mean Recall at rank 1 (MS R1, %) averaged across image-to-text and text-to-image on the COCO 5K test set:

    Prediction Mode Dynamic Partitioning Progressive Distillation ImageNet Top-1 (%) MS COCO R1 (%)
    Baseline (No Distillation) - - 8.18 27.58
    Forward Bootstrapping No No 6.23 20.40
    Swapped Prediction No No 8.46 23.51
    Forward Bootstrapping Yes No 8.37 25.64
    Swapped Prediction Yes No 8.81 26.24
    Forward Bootstrapping Yes Yes 8.87 26.18
    Swapped Prediction Yes Yes 9.51 28.48

    Key observations:

    1. Directly applying standard forward bootstrapping with static batch partitioning causes significant performance degradation (ImageNet Top-1 drops from 8.18% to 6.23%, MS R1 drops from 27.58% to 20.40%) due to error reinforcement and partial representation collapse.
    2. Swapped prediction mitigates collapse and outperforms forward bootstrapping under all tested conditions.
    3. Combining swapped prediction, dynamic batch partitioning, and progressive cosine annealing achieves the highest zero-shot accuracy (9.51%) and retrieval recall (28.48%).
  9. Knowl 9 — Effective Robustness Under Natural Distribution Shifts

    empirical result

    When evaluated on out-of-distribution datasets (ImageNet-A, ImageNet-O, ImageNet-R, and ImageNetV2) against in-distribution ImageNet Top-1 accuracy, models trained with Progressive Self-Distillation demonstrate higher effective robustness than both standard ImageNet-trained models and CLIP.

    Linear fits relating ImageNet in-distribution accuracy (xx) to mean transfer accuracy (yy) across the four out-of-distribution test sets are:

    • ImageNet-trained models: y=0.486x+0.52y = 0.486x + 0.52
    • CLIP models: y=0.808x+1.54y = 0.808x + 1.54
    • Progressive Self-Distillation (Ours): y=0.914x+0.44y = 0.914x + 0.44

    The steeper slope (0.9140.914 vs. 0.8080.808) demonstrates that as the base visual representation quality scales, the out-of-distribution robustness advantage of progressive self-distillation over CLIP widens.

  10. Knowl 10 — Representation Embedding Similarity Distribution

    empirical result

    Comparing cross-modal cosine similarity distributions between paired (positive) and unpaired (negative) image-text samples on the MS COCO test set reveals key structural differences in the learned representation geometry:

    1. Positive Pairs: Progressive self-distillation produces cosine similarity distributions with a higher mean and lower variance than baseline CLIP and OpenAI's pretrained CLIP.
    2. Negative Pairs: While baseline CLIP forces negative sample similarity scores to concentrate tightly around zero, progressive self-distillation shifts the negative similarity distribution toward higher values, accommodating non-zero semantic overlap among batch negatives.

    Allowing soft alignment among semantically related negative instances reduces conflicting repulsive gradient updates, resulting in a more globally consistent joint embedding space.

Coverage note — No substantial contributed material was omitted; the extracted knowls cover the loss objective formulation, swapped prediction target mechanism, progressive scheduling and dynamic partitioning, experimental configuration, zero-shot classification, linear probe evaluation, retrieval performance, component ablations, effective robustness analysis, and representation similarity distributions.

References

  1. 1.Görkem Algan and Ilkay Ulusoy. Metalabelnet: Learning to generate soft-labels from noisy-labels. arXiv preprint arXiv:2103.10869, 2021. 1, 2
  2. 2.Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi. A theoretical analysis of contrastive unsupervised representation learning. arXiv preprint arXiv:1902.09229, 2019. 2
  3. 3.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. 5, 6
  4. 4.Thomas Berg, Jiongxin Liu, Seung Woo Lee, Michelle L. Alexander, David W. Jacobs, and Peter N. Belhumeur. Birdsnap: Large-scale fine-grained visual categorization of birds. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), June 2014. 7
  5. 5.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In European Conference on Computer Vision, 2014. 7
  6. 6.Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 535–541, 2006. 2
  7. 7.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the European Conference on Computer Vision (ECCV), pages 132–149, 2018. 2
  8. 8.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882, 2020. 2
  9. 9.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3558–3568, 2021. 1, 2, 5, 6
  10. 10.Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. Learning efficient object detection models with knowledge distillation. Advances in neural information processing systems, 30, 2017. 2
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML), 2020. 2
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Ieee, 2009. 5, 6, 7
  13. 13.Aditya Deshpande, Jason Rock, and David Forsyth. Learning large-scale automatic image colorization. In Proceedings of the IEEE International Conference on Computer Vision, pages 567–575, 2015. 2
  14. 14.Qianggang Ding, Sifan Wu, Hao Sun, Jiadong Guo, and Shu-Tao Xia. Adaptive regularization of labels. arXiv preprint arXiv:1908.05474, 2019. 2
  15. 15.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE international conference on computer vision, pages 1422–1430, 2015. 2
  16. 16.Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. Discriminative unsupervised feature learning with exemplar convolutional neural networks. IEEE transactions on pattern analysis and machine intelligence, 38(9):1734–1747, 2015. 2
  17. 17.Alexey Dosovitskiy, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. Discriminative unsupervised feature learning with convolutional neural networks. Advances in neural information processing systems, 27:766–774, 2014. 2
  18. 18.Li Fei-Fei, Rob Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006. 6
  19. 19.Andreas Fürst, Elisabeth Rumetshofer, Viet Tran, Hubert Ramsauer, Fei Tang, Johannes Lehner, David Kreil, Michael Kopp, Günter Klambauer, Angela Bitto-Nemling, et al. Cloob: Modern hopfield networks with infoloob outperform clip. arXiv preprint arXiv:2110.11316, 2021. 1, 2
  20. 20.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728, 2018. 2
  21. 21.Sangchul Hahn and Heeyoul Choi. Self-knowledge distillation in natural language processing. arXiv preprint arXiv:1908.01851, 2019. 2
  22. 22.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. Co-teaching: Robust training of deep neural networks with extremely noisy labels. arXiv preprint arXiv:1804.06872, 2018. 2
  23. 23.Tengda Han, Weidi Xie, and Andrew Zisserman. Video representation learning by dense predictive coding. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 2
  24. 24.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2, 4
  25. 25.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021. 1, 5, 6
  26. 26.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples.(2019). arXiv preprint cs.LG/1907.07174, 2019. 5, 6
  27. 27.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 2, 4
  28. 28.Gabriel Ilharco, Mitchell Wortsman, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip, July 2021. If you use this software, please cite it as below. 5
  29. 29.J Yu Jason, Adam W Harley, and Konstantinos G Derpanis. Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness. In European Conference on Computer Vision, pages 3–10. Springer, 2016. 2
  30. 30.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv preprint arXiv:2102.05918, 2021. 1
  31. 31.Longlong Jing and Yingli Tian. Self-supervised visual feature learning with deep neural networks: A survey. IEEE transactions on pattern analysis and machine intelligence, 2020. 2
  32. 32.Dahun Kim, Donghyeon Cho, and In So Kweon. Self-supervised video representation learning with space-time cubic puzzles. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 8545–8552, 2019. 2
  33. 33.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  34. 34.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 6
  35. 35.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Learning representations for automatic colorization. In European conference on computer vision, pages 577–593. Springer, 2016. 2
  36. 36.Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang. Unsupervised representation learning by sorting sequences. In Proceedings of the IEEE International Conference on Computer Vision, pages 667–676, 2017. 2
  37. 37.Junnan Li, Yongkang Wong, Qi Zhao, and Mohan S Kankanhalli. Learning to learn from noisy labeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5051–5059, 2019. 2
  38. 38.Tianhong Li, Jianguo Li, Zhuang Liu, and Changshui Zhang. Few sample knowledge distillation for efficient network compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14639–14647, 2020. 2, 4
  39. 39.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. arXiv preprint arXiv:2110.05208, 2021. 1, 2
  40. 40.Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li. Learning from noisy labels with distillation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1910–1918, 2017. 2
  41. 41.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 5
  42. 42.Yuanze Lin, Xun Guo, and Yan Lu. Self-supervised video representation learning with meta-contrastive network. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8239–8249, 2021. 4
  43. 43.Haoliang Liu, Tan Yu, and Ping Li. Inflate and shrink: Enriching and reducing interactions for fast text-image retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9796–9809, 2021. 2
  44. 44.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 5
  45. 45.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017. 5
  46. 46.John P Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt. Accuracy on the line: On the strong correlation between out-of-distribution and in-distribution generalization. In International Conference on Machine Learning, pages 7721–7735. PMLR, 2021. 6
  47. 47.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6707–6717, 2020. 2
  48. 48.Ishan Misra, C Lawrence Zitnick, and Martial Hebert. Shuffle and learn: unsupervised learning using temporal order verification. In European Conference on Computer Vision, pages 527–544. Springer, 2016. 2
  49. 49.Hossein Mobahi, Ronan Collobert, and Jason Weston. Deep learning from temporal coherence in video. In Proceedings of the 26th Annual International Conference on Machine Learning, pages 737–744, 2009. 2
  50. 50.Pedro Morgado, Ishan Misra, and Nuno Vasconcelos. Robust audio-visual instance discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12934–12945, 2021. 2, 3
  51. 51.Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. Audio-visual instance discrimination with cross-modal agreement. CoRR, abs/2004.12943, 2020. 2
  52. 52.Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. Audio-visual instance discrimination with cross-modal agreement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12475–12486, 2021. 1
  53. 53.Rafael Müller, Simon Kornblith, and Geoffrey Hinton. When does label smoothing help? arXiv preprint arXiv:1906.02629, 2019. 2
  54. 54.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision, pages 69–84. Springer, 2016. 2
  55. 55.Curtis Northcutt, Lu Jiang, and Isaac Chuang. Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research, 70:1373–1411, 2021. 2
  56. 56.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 3
  57. 57.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012. 7
  58. 58.Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan. Learning features by watching objects move. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2701–2710, 2017. 2
  59. 59.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2536–2544, 2016. 2
  60. 60.Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu. Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1944–1952, 2017. 2
  61. 61.Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton. Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548, 2017. 2
  62. 62.AJ Piergiovanni, Anelia Angelova, and Michael S Ryoo. Evolving losses for unsupervised video representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 133–142, 2020. 2
  63. 63.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021. 1, 2, 3, 5, 6
  64. 64.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In International Conference on Machine Learning, pages 5389–5400. PMLR, 2019. 5, 6
  65. 65.Scott Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich. Training deep neural networks on noisy labels with bootstrapping. arXiv preprint arXiv:1412.6596, 2014. 2, 3
  66. 66.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014. 2
  67. 67.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018. 1, 2, 5, 6
  68. 68.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021. 1, 2
  69. 69.Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. Meta-weight-net: Learning an explicit mapping for sample weighting. arXiv preprint arXiv:1902.07379, 2019. 2
  70. 70.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 1, 2, 6
  71. 71.Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. 1, 2
  72. 72.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In European Conference on Computer Vision (ECCV), 2020. 2
  73. 73.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, pages 776–794. Springer, 2020. 2
  74. 74.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pages 1096–1103, 2008. 2
  75. 75.Jue Wang, Haofan Wang, Jincan Deng, Weijia Wu, and Debing Zhang. Efficientclip: Efficient cross-modal pre-training by ensemble confident learning and language modeling. arXiv preprint arXiv:2109.04699, 2021. 1, 2
  76. 76.Yisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo, Jinfeng Yi, and James Bailey. Symmetric cross entropy for robust learning with noisy labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 322–330, 2019. 2
  77. 77.Zhirong Wu, Yuanjun Xiong, X Yu Stella, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 2
  78. 78.Ting-Bing Xu and Cheng-Lin Liu. Data-distortion guided self-distillation for deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 5565–5572, 2019. 2
  79. 79.Zhichao Yin and Jianping Shi. Geonet: Unsupervised learning of dense depth, optical flow and camera pose. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1983–1992, 2018. 2
  80. 80.Sukmin Yun, Jongjin Park, Kimin Lee, and Jinwoo Shin. Regularizing class-wise predictions via self-knowledge distillation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13876–13885, 2020. 2
  81. 81.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In European conference on computer vision, pages 649–666. Springer, 2016. 2
  82. 82.Richard Zhang, Phillip Isola, and Alexei A Efros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1058–1067, 2017. 2
  83. 83.Zhilu Zhang and Mert R Sabuncu. Generalized cross entropy loss for training deep neural networks with noisy labels. In 32nd Conference on Neural Information Processing Systems (NeurIPS), 2018. 2
  84. 84.Zizhao Zhang, Han Zhang, Sercan O Arik, Honglak Lee, and Tomas Pfister. Distilling effective supervision from severe label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9294–9303, 2020. 2
  85. 85.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017. 5, 6
  86. 86.Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, and Daniel Yamins. Unsupervised learning from video with deep neural embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9563–9572, 2020. 2
  87. 87.Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins. Local aggregation for unsupervised learning of visual embeddings. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6002–6012, 2019. 2
  88. 88.Mohammadreza Zolfaghari, Yi Zhu, Peter Gehler, and Thomas Brox. Crossclr: Cross-modal contrastive learning for multi-modal video representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1450–1459, 2021. 2

Citation

MLA
Andonian, A., et al. “Robust Cross-Modal Representation Learning with Progressive Self-Distillation”. arXiv, 2022, http://arxiv.org/abs/2204.04588v1.
APA
Andonian, A., Chen, S., & Hamid, R. (2022). Robust Cross-Modal Representation Learning with Progressive Self-Distillation. arXiv. http://arxiv.org/abs/2204.04588v1
Chicago
Andonian, A., S. Chen, and R. Hamid. 2022. “Robust Cross-Modal Representation Learning with Progressive Self-Distillation”. arXiv. http://arxiv.org/abs/2204.04588v1.
Harvard
Andonian, A., Chen, S. and Hamid, R. (2022) “Robust Cross-Modal Representation Learning with Progressive Self-Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.04588v1.
Vancouver
1. Andonian A, Chen S, Hamid R (2022) Robust Cross-Modal Representation Learning with Progressive Self-Distillation. arXiv

BibTeX

@article{andonian2022robust,
  title = {Robust Cross-Modal Representation Learning with Progressive Self-Distillation},
  author = {Andonian, Alex and Chen, Shixing and Hamid, Raffay},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.04588v1},
  eprint = {2204.04588}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE