Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Alexei BaevskiArun BabuWei-Ning HsuMichael Auli

article2023ICML162 citations

Presents data2vec 2.0, a unified self-supervised learning algorithm across vision, speech, and text that achieves up to a 16-fold pre-training speedup over existing models like MAE and wav2vec 2.0 while maintaining competitive performance.

Listen

Self-supervised machine learning allows models to learn rich data representations without manual labeling, but the initial training phase demands enormous computational power, long timelines, and massive hardware clusters. Furthermore, most existing algorithms are designed for single data types, preventing organizations from using a single, unified training pipeline across modalities. The article addresses these efficiency and scalability bottlenecks by evaluating data2vec 2.0, an optimized self-supervised training algorithm designed to dramatically cut computational costs and training time while generalizing across computer vision, speech recognition, and natural language processing.

The article demonstrates this improved framework by conducting extensive empirical pre-training and downstream evaluation benchmarks across standard datasets: ImageNet-1K for vision, Librispeech and Libri-light for speech, and the GLUE benchmark for language. The core approach pairs a teacher model that analyzes full input samples to produce contextual target representations with a student model that learns to predict these targets from partially masked inputs. To maximize efficiency, the authors introduce three primary architectural optimizations: skipping the encoding of masked inputs, adopting a fast, lightweight convolutional decoder, and amortizing teacher overhead by reusing a single target representation across multiple masked versions of each sample.

The key findings reveal major training speedups across all three modalities without sacrificing downstream performance. For computer vision, the algorithm matches the accuracy of Masked Autoencoders with a 16.4-fold reduction in pre-training time (just over 3 hours versus 50.7 hours) while running only 20 training epochs instead of 1,600. For speech recognition, it matches wav2vec 2.0 performance in 10.6-fold less time, while reducing relative word error rates by up to 26% on low-resource benchmarks. For natural language processing, it matches retrained RoBERTa benchmarks in roughly half the wall-clock time and nearly an eightfold reduction in epochs. Additionally, the multi-mask strategy enables robust pre-training with significantly smaller batch sizes (such as 512 images instead of 4,096).

These results demonstrate that creating rich, contextualized learning targets significantly accelerates the learning process rather than slowing it down. In practical terms, organizations can achieve state-of-the-art representation quality with substantially lower cloud computing budgets, reduced energy consumption, and faster experimentation cycles. Because the algorithm maintains a unified objective across vision, speech, and text, it also simplifies operational workflows by eliminating the need to maintain distinct training architectures for each data modality.

Decision-makers should consider adopting this architecture when training foundation models or updating self-supervised pipelines, especially where compute budgets or hardware availability are constrained. While the findings provide high confidence across standard benchmarks, the evaluations remain limited to unimodal models trained independently for each modality and do not yet cover joint multimodal representations (such as simultaneous vision-language models) or modalities beyond image, speech, and text. Future initiatives should focus on piloting the framework in broader production settings and expanding it into multimodal and video domains.

Cover for Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Abstract

Current self-supervised learning algorithms are often modality-specific and require large amounts of computational resources. To address these issues, we increase the training efficiency of data2vec, a learning objective that generalizes across several modalities. We do not encode masked tokens, use a fast convolutional decoder and amortize the effort to build teacher representations. data2vec 2.0 benefits from the rich contextualized target representations introduced in data2vec which enable a fast self-supervised learner. Experiments on ImageNet-1K image classification show that data2vec 2.0 matches the accuracy of Masked Autoencoders in 16.4x lower pre-training time, on Librispeech speech recognition it performs as well as wav2vec 2.0 in 10.6x less time, and on GLUE natural language understanding it matches a retrained RoBERTa model in half the time. Trading some speed for accuracy results in ImageNet-1K top-1 accuracy of 86.8% with a ViT-L model trained for 150 epochs. Models and code are available at www.github.com/pytorch/fairseq/tree/master/examples/data2vec.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Contextualized Target Prediction
  • 3.2. Model Architecture
  • 3.3. Multi-mask Training
  • 3.4. Inverse Block Masking
  • 4. Experiments
  • 4.1. Efficiency
  • 4.2. Computer Vision
  • 4.3. Speech Processing
  • 4.4. Natural Language Processing
  • 4.5. Ablations
  • 5. Conclusion and Future Work
  • References
  • A. Pre-training Hyper-parameters
  • B. Effect of Pre-training Dataset Size

Knowls

  1. Knowl 1 — Shared contextualized-target prediction objective

    model/method

    data2vec 2.0 uses the same teacher–student learning objective for vision, speech, and language, while training a separate model for each modality. A Transformer teacher processes the complete, unmasked example, so its representations are contextualized by information from the whole example. At each position, the target is formed by instance-normalizing activations from the teacher’s top KK feed-forward blocks and averaging them. A student predicts these targets from a masked version of the example using an L2L_2 regression loss; for images, the experiments also use an optional global CLS-token loss.

    The teacher parameters Δ\Delta track the student encoder parameters θ\theta by exponential moving average:

    Δ←τΔ+(1−τ)θ.\Delta \leftarrow \tau\Delta + (1-\tau)\theta.

    Here τ\tau is a scalar averaging coefficient. It increases linearly from an initial value τ0\tau_0 to a final value τe\tau_e over τn\tau_n training updates, then remains at τe\tau_e. The teacher target is computed from the unmasked sample and reused for the student’s masked prediction task.

  2. Knowl 2 — Asymmetric sparse-encoding architecture with convolutional decoder

    model/method

    The student encoder processes only unmasked patches or time steps, rather than encoding the entire masked sequence. Its outputs are restored to their original positions alongside fixed representations for masked positions; the experiments found random Gaussian noise sufficient for those masked-position representations. A lightweight decoder then predicts teacher targets at masked positions. This asymmetric design avoids much of the student-side computation associated with masked inputs while retaining full-sample teacher targets.

    The decoder has DD convolutional layers, each followed by layer normalization and GELU and equipped with a residual connection. It uses 1-D convolutions for sequential data and 2-D convolutions for images, with grouped convolutions for efficiency. Modality-specific input encoders map image patches, speech features, or byte-pair-encoded text into a Transformer representation.

  3. Knowl 3 — Multi-mask training amortizes teacher computation

    model/method

    For each unmasked training example, data2vec 2.0 computes the teacher target once and trains on MM different masked versions of that same example against the shared target. It also reuses the feature-encoder output across those versions, avoiding repeated feature extraction; this is especially useful for speech, where feature encoding is a substantial part of the computation. Increasing MM spreads the cost of full-sample teacher computation over more student predictions and allows effective training with smaller batches.

    In a reduced ImageNet-1K experiment with 100,000 updates, using M=16M=16 rather than M=2M=2 at batch size 64 raised top-1 accuracy by 4.6 percentage points with other settings held fixed. The batch-size-64 run with M=1M=1 diverged, illustrating that multiple masks can also provide enough learning signal for a small batch.

  4. Knowl 4 — Inverse block masking preserves contiguous context

    algorithm

    Inverse block masking selects contiguous regions to preserve, rather than selecting regions to mask. For a sample with LL positions, mask ratio RR, block parameter BB, and mask-adjustment parameter AA, the reported rule sets the number of block starting points to L((1−R)+A)/BL((1-R)+A)/B. Starting points are sampled without replacement. Each selected region is expanded symmetrically: to width BB for speech and text, or to a square with side length B\sqrt{B} for images. Regions may overlap, so the resulting mask can initially leave fewer positions unmasked than intended. The procedure then randomly changes individual positions until each sample has the target of L(1−R)L(1-R) unmasked positions.

    The contiguous preserved regions are intended to give the student local context for its predictions while still allowing masked positions to be omitted from student encoding. The authors report that 0.05<A<0.150.05<A<0.15 worked well. In the masking ablation, B=1B=1 corresponds to random masking.

  5. Knowl 5 — ImageNet-1K accuracy with substantially less pre-training

    empirical result

    After self-supervised pre-training on unlabeled ImageNet-1K and fine-tuning for image classification, data2vec 2.0 achieved higher top-1 validation accuracy than MAE with fewer training epochs. The principal comparisons were:

    • ViT-B: data2vec 2.0 reached 84.5% after 200 epochs; MAE reached 83.6% after 1,600 epochs. The reported pre-training times were 32 hours and 50.7 hours, respectively, on 32 A100 GPUs.
    • ViT-L: data2vec 2.0 reached 86.8% after 150 epochs; MAE reached 85.9% after 1,600 epochs. The reported times were 63.3 hours and 93.3 hours, respectively, on 32 A100 GPUs.
    • ViT-H/14: data2vec 2.0 reached 87.4% after 100 epochs and 66.1 hours; MAE reached 86.9% after 1,600 epochs and 113.6 hours, measured on 64 A100 GPUs.

    At a shorter-training operating point, data2vec 2.0 reached 83.7% in just over 3 hours, compared with MAE’s 83.6% in 50.7 hours, a reported 16.4-fold speed-up. Thus the method offers different speed–accuracy trade-offs, from rapid matching of MAE to higher final accuracy with longer training.

  6. Knowl 6 — Speech recognition results on LibriSpeech and Libri-light

    empirical result

    The models were fine-tuned for speech recognition and evaluated on the LibriSpeech test-other set with 4-gram language models. Word error rates (WER, lower is better) are listed in order of labeled fine-tuning data: 10 minutes, 1 hour, 10 hours, 100 hours, and 960 hours. Base models were pre-trained on 960 hours of unlabeled LibriSpeech audio; Large models were pre-trained on 60,000 hours of Libri-light audio.

    • Base wav2vec 2.0: 15.6, 11.3, 9.5, 8.0, 6.1 WER; pre-training took 57.3 hours.
    • Base data2vec: 12.3, 9.1, 8.1, 6.8, 5.5 WER; 63.3 hours.
    • Base data2vec 2.0: 11.5, 8.7, 7.6, 6.4, 5.2 WER; 43.3 hours.
    • Large wav2vec 2.0: 10.3, 7.1, 5.8, 4.6, 3.6 WER; 150.0 hours.
    • Large data2vec: 8.4, 6.3, 5.3, 4.6, 3.7 WER; 108.0 hours.
    • Large data2vec 2.0: 8.4, 6.3, 5.1, 4.3, 3.5 WER; 76.7 hours.

    The Base and Large data2vec 2.0 runs used 16 and 64 A100 GPUs, respectively. Against wav2vec 2.0, the paper reports relative WER reductions of up to 26% for Base and 18% for Large models, while using less pre-training time.

  7. Knowl 7 — GLUE performance with a shorter language-model pre-training run

    empirical result

    For single-task fine-tuning on GLUE, data2vec 2.0 achieved an average score of 82.6 after 4.1 epochs and 28.2 hours of pre-training on 16 A100 GPUs. A reproduced RoBERTa/BERT-setup baseline scored 82.5 after 31.8 epochs and 50.5 hours; data2vec scored 82.7 after 31.8 epochs and 69.4 hours. The data2vec 2.0 result therefore matched the reproduced baseline’s average while using about half its pre-training time, and was 2.5 times faster than data2vec.

    The average uses these development-set metrics: accuracy for MNLI, QNLI, RTE, QQP, and SST-2; the unweighted mean of accuracy and F1 for MRPC; the unweighted mean of Pearson and Spearman correlation for STS-B; and Matthews correlation for CoLA. The task scores for the reproduced baseline and data2vec 2.0, respectively, were MNLI matched/unmatched 84.1/83.9 and 83.7/83.7; QNLI 90.4 and 90.7; RTE 69.3 and 68.6; MRPC 89.0 and 88.8; QQP 89.3 and 89.3; STS-B 88.9 and 87.3; CoLA 56.8 and 59.1; and SST-2 92.3 and 92.9. The data2vec 2.0 language model used a 42% masking rate, compared with 15% for the BERT-style baseline.

  8. Knowl 8 — Vision ablations support contextual targets and inverse block masking

    empirical result

    ImageNet-1K ablations reported top-1 accuracy for data2vec 2.0 vision models. In the training-loss ablation, the baseline scored 84.4%; removing the global CLS loss scored 84.2%; adding pixel regression to contextualized-target prediction scored 84.3%; and training with pixel regression alone scored 83.5%. These results indicate that pixel regression did not improve on contextualized-target prediction alone, while removing the CLS loss caused a small reduction.

    In the masking ablation, block masking with B=3B=3 scored 84.1%. Inverse block masking scored 83.7% at B=1B=1, 84.4% at B=2B=2, 84.4% at B=3B=3, and 84.2% at B=4B=4. Thus the tested inverse-block settings with B=2B=2 or B=3B=3 outperformed random masking (B=1B=1) and block masking at B=3B=3.

  9. Knowl 9 — Speech-specific learned ALiBi scaling improves recognition

    model/method

    For speech, data2vec 2.0 adds a distance-proportional bias to Transformer query–key attention scores. The bias values are initialized as in ALiBi and kept frozen; training learns a scalar for each attention head, initialized to 1.0. This modifies the speech setup while adding few parameters; the paper reports 16 such parameters for a Large model.

    On dev-other after pre-training on LibriSpeech-960 and fine-tuning with the 10-hour Libri-light labeled set, the baseline WER was 10.9. Removing ALiBi increased WER to 11.3; using one shared scale across heads instead of learned per-head scales gave 11.0; and not learning the scales gave 11.7. These ablations support both the attention bias and the learned scaling design under this evaluation setup.

  10. Knowl 10 — Larger vision models benefit more from additional pre-training data

    empirical result

    The authors randomly subsampled ImageNet-1K pre-training data at fractions ranging from 1% to 100%, while holding the pre-training hyperparameters fixed and fine-tuning every model on the full ImageNet-1K dataset. Increasing the pre-training data improved the ViT-L model more than the ViT-B model. The authors interpret this pattern as suggesting that the Base model underfits ImageNet-1K with the data2vec-style pre-training task; the result is specific to the tested models, dataset, and setup.

Coverage note — The full modality-specific pre-training hyperparameter appendices and detailed results for comparison methods are not reproduced as separate knowls; they are configuration or baseline details rather than standalone contributions.

References

  1. 1.Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text, 2021.
  2. 2.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K. Flamingo: a visual language model for few-shot learning. arXiv, abs/2204.14198, 2022.
  3. 3.Assran, M., Caron, M., Misra, I., Bojanowski, P., Bordes, F., Vincent, P., Joulin, A., Rabbat, M., and Ballas, N. Masked siamese networks for label-efficient learning. arXiv, abs/2204.07141, 2022.
  4. 4.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv, abs/1607.06450, 2016.
  5. 5.Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., Baevski, A., Conneau, A., and Auli, M. Xls-r: Self-supervised cross-lingual speech representation learning at scale. In Proc. of Interspeech, 2022.
  6. 6.Baevski, A., Edunov, S., Liu, Y., Zettlemoyer, L., and Auli, M. Cloze-driven pretraining of self-attention networks. In Proc. of EMNLP, 2019.
  7. 7.Baevski, A., Schneider, S., and Auli, M. vq-wav2vec: Self-supervised learning of discrete speech representations. In Proc. of ICLR, 2020a.
  8. 8.Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proc. of NeurIPS, 2020b.
  9. 9.Baevski, A., Hsu, W.-N., Conneau, A., and Auli, M. Unsupervised speech recognition. In Proc. of NeurIPS, 2021.
  10. 10.Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M. data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv, abs/2202.03555, 2022.
  11. 11.Bao, H., Dong, L., and Wei, F. Beit: BERT pre-training of image transformers. arXiv, abs/2106.08254, 2021.
  12. 12.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Proc. of NeurIPS, 2020.
  13. 13.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. arXiv, abs/2006.09882, 2020a.
  14. 14.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. arXiv, abs/2006.09882, 2020b.
  15. 15.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. arXiv, abs/2104.14294, 2021.
  16. 16.Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y., Qian, Y., Wu, J., Zeng, M., and Wei, F. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. arXiv, abs/2110.13900, 2021a.
  17. 17.Chen, X., Xie, S., and He, K. An empirical study of training self-supervised vision transformers. arXiv, abs/2104.02057, 2021b.
  18. 18.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, L., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, A., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways. arXiv, abs/2204.02311, 2022.
  19. 19.Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR, 2020. URL https://openreview.net/pdf?id=r1xMH1BtvB.
  20. 20.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. Proc. of NAACL, 2019.
  21. 21.Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N. Peco: Perceptual codebook for bert pre-training of vision transformers, 2022.
  22. 22.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv, abs/2010.11929, 2020.
  23. 23.Eloff, R., Nortje, A., van Niekerk, B., Govender, A., Nortje, L., Pretorius, A., Van Biljon, E., van der Westhuizen, E., van Staden, L., and Kamper, H. Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks. arXiv, abs/1904.07556, 2019.
  24. 24.Gao, S., Zhou, P., Cheng, M.-M., and Yan, S. Towards sustainable self-supervised learning. arXiv, abs/2210.11016, 2022.
  25. 25.Girdhar, R., El-Nouby, A., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. Omnimae: Single model masked pre-training on images and videos. arXiv, 2022.
  26. 26.Grill, J.-B., Strub, F., Altche, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. Bootstrap your own latent: A new approach to self-supervised learning. arXiv, abs/2006.07733, 2020.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Deep Residual Learning for Image Recognition. In Proc. of CVPR, 2015.
  28. 28.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. arXiv, abs/2111.06377, 2021.
  29. 29.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv, abs/1606.08415, 2016.
  30. 30.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L. Training compute-optimal large language models. arXiv, 2022.
  31. 31.Hsu, W.-N., Tsai, Y.-H. H., Bolte, B., Salakhutdinov, R., and Mohamed, A. Hubert: How much can a bad teacher benefit ASR pre-training? In Proc. of ICASSP, 2021.
  32. 32.Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O., Botvinick, M. M., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver io: A general architecture for structured inputs & outputs. arXiv, abs/2107.14795, 2021a.
  33. 33.Jaegle, A., Gimeno, F., Brock, A., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver: General perception with iterative attention. arXiv, abs/2103.03206, 2021b.
  34. 34.Jing, L., Zhu, J., and LeCun, Y. Masked siamese convnets. arXiv, abs/2206.07700, 2022.
  35. 35.Kahn, J., Riviere, M., Zheng, W., Kharitonov, E., Xu, Q., Mazare, P., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E. Libri-light: A benchmark for asr with limited or no supervision. arXiv, abs/1912.07875, 2019.
  36. 36.Kahn, J. et al. Libri-light: A benchmark for asr with limited or no supervision. In Proc. of ICASSP, 2020.
  37. 37.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Proc. of NIPS, 2012.
  38. 38.Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. arXiv, abs/1909.11942, 2019.
  39. 39.Li, C., Yang, J., Zhang, P., Gao, M., Xiao, B., Dai, X., Yuan, L., and Gao, J. Efficient self-supervised vision transformers for representation learning. arXiv, abs/2106.09785, 2021.
  40. 40.Liu, A. T., Li, S.-W., and Lee, H.-y. Tera: Self-supervised learning of transformer encoder representation for speech. IEEE/ACM Trans. on Audio, Speech, and Language Processing, 2021.
  41. 41.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  42. 42.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In Proc. of NAACL System Demonstrations, 2019.
  43. 43.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an asr corpus based on public domain audio books. In Proc. of ICASSP, pp. 5206–5210. IEEE, 2015.
  44. 44.Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization. arXiv, abs/1705.04304, 2017.
  45. 45.Peng, Z., Dong, L., Bao, H., Ye, Q., and Wei, F. Beit v2: Masked image modeling with vector-quantized visual tokenizers. arXiv, abs/2208.06366, 2022.
  46. 46.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv, abs/2108.12409, 2021.
  47. 47.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf, 2018.
  48. 48.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv, abs/2103.00020, 2021.
  49. 49.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv, abs/1910.10683, 2019.
  50. 50.Schneider, S., Baevski, A., Collobert, R., and Auli, M. wav2vec: Unsupervised pre-training for speech recognition. In Proc. of Interspeech, 2019.
  51. 51.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proc. of ACL, 2016.
  52. 52.Shi, B., Mohamed, A., and Hsu, W.-N. Learning lip-based audio-visual speaker embeddings with av-hubert. arXiv, abs/2205.07180, 2022.
  53. 53.Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. FLAVA: A foundational language and vision alignment model. arXiv, abs/2112.04482, 2021.
  54. 54.Ulyanov, D., Vedaldi, A., and Lempitsky, V. S. Instance normalization: The missing ingredient for fast stylization. arXiv, abs/1607.08022, 2016.
  55. 55.van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. Proc. of NIPS, 2018.
  56. 56.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Proc. of NIPS, 2017.
  57. 57.Vyas, A., Hsu, W.-N., Auli, M., and Baevski, A. On-demand compute reduction with stochastic wav2vec 2.0. In Proc. of Interspeech, 2022.
  58. 58.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv, abs/1804.07461, 2018.
  59. 59.Wang, W., Bao, H., Dong, L., and Wei, F. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv, abs/2111.02358, 2021.
  60. 60.Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. arXiv, abs/2112.09133, 2021.
  61. 61.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent abilities of large language models. arXiv, abs/2206.07682, 2022.
  62. 62.Wu, F., Kim, K., Pan, J., Han, K. J., Weinberger, K. Q., and Artzi, Y. Performance-efficiency trade-offs in unsupervised pre-training for speech recognition. In Proc. of Interspeech, 2022a.
  63. 63.Wu, Z., Wang, S., Gu, J., Khabsa, M., Sun, F., and Ma, H. CLEAR: contrastive learning for sentence representation. arXiv, abs/2012.15466, 2020.
  64. 64.Wu, Z., Lai, Z., Sun, X., and Lin, S. Extreme masking for learning instance and distributed visual representations. arXiv, 2206.04667, 2022b.
  65. 65.Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. Simmim: A simple framework for masked image modeling. arXiv, abs/2111.09886, 2021.
  66. 66.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. ibot: Image bert pre-training with online tokenizer, 2021.
  67. 67.Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. arXiv, abs/1506.06724, 2015.

Citation

MLA
Baevski, A., et al. “Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language”. International Conference on Machine Learning, vol. 202, 2023, pp. 1416–29, https://proceedings.mlr.press/v202/baevski23a.html.
APA
Baevski, A., Babu, A., Hsu, W.-N., & Auli, M. (2023). Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language. International Conference on Machine Learning, 202, 1416–1429. https://proceedings.mlr.press/v202/baevski23a.html
Chicago
Baevski, A., A. Babu, W.-N. Hsu, and M. Auli. 2023. “Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language”. International Conference on Machine Learning 202: 1416–29. https://proceedings.mlr.press/v202/baevski23a.html.
Harvard
Baevski, A. et al. (2023) “Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language”, International Conference on Machine Learning. PMLR, pp. 1416–1429. Available at: https://proceedings.mlr.press/v202/baevski23a.html.
Vancouver
1. Baevski A, Babu A, Hsu W-N, Auli M (2023) Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language. In: International Conference on Machine Learning. PMLR, pp 1416–1429

BibTeX

@InProceedings{pmlr-v202-baevski23a,
  title = 	 {Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language},
  author =       {Baevski, Alexei and Babu, Arun and Hsu, Wei-Ning and Auli, Michael},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {1416--1429},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/baevski23a/baevski23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/baevski23a.html},
  abstract = 	 {Current self-supervised learning algorithms are often modality-specific and require large amounts of computational resources. To address these issues, we increase the training efficiency of data2vec, a learning objective that generalizes across several modalities. We do not encode masked tokens, use a fast convolutional decoder and amortize the effort to build teacher representations. data2vec 2.0 benefits from the rich contextualized target representations introduced in data2vec which enable a fast self-supervised learner. Experiments on ImageNet-1K image classification show that data2vec 2.0 matches the accuracy of Masked Autoencoders in 16.4x lower pre-training time, on Librispeech speech recognition it performs as well as wav2vec 2.0 in 10.6x less time, and on GLUE natural language understanding it matches a retrained RoBERTa model in half the time. Trading some speed for accuracy results in ImageNet-1K top-1 accuracy of 86.8% with a ViT-L model trained for 150 epochs.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/