data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Alexei BaevskiWei-Ning HsuQiantong XuArun BabuJiatao GuMichael Auli

article2022ICML1,259 citations

Introduces data2vec, a unified self-supervised learning framework that predicts contextualized latent representations from masked inputs across speech, vision, and language, matching or outperforming modality-specific approaches.

Listen

Self-supervised learning enables artificial intelligence systems to learn rich representations from raw data without human-annotated labels. However, current algorithms are fragmented: speech, computer vision, and natural language processing each rely on distinct, modality-specific objectives, such as predicting discrete words, visual tokens, or quantized speech units. This fragmentation complicates development and prevents machine learning models from using a unified learning principle across different domains.

The article introduces and evaluates data2vec, a general self-supervised framework designed to apply a single, unified learning objective across speech, vision, and text. The approach uses a standard Transformer architecture operated in two modes: a teacher network and a student network. The teacher encodes an unmasked input to generate continuous, contextualized internal representations, while the student receives a masked version of the same input and learns to predict the teacher's representations. The teacher's parameters are updated over time as an exponentially moving average of the student's parameters. To provide richer training targets, the model averages representations across multiple network layers rather than predicting only the final layer or local raw features.

Empirical evaluations demonstrate that data2vec matches or outperforms leading domain-specific systems across all three modalities. In computer vision, it achieved top-1 accuracy on ImageNet-1K of 84.2% with a standard Vision Transformer Base (ViT-B) and 86.6% with Large (ViT-L), setting a new state of the art among single self-supervised models. In speech processing on LibriSpeech, data2vec reduced word error rates compared to established methods like wav2vec 2.0 and HuBERT, achieving a 20% relative word error rate reduction in low-resource settings with only 10 minutes of labeled data. For audio event classification on AudioSet, it reached a leading 34.5 mean average precision. In natural language processing, data2vec matched or slightly exceeded a comparable RoBERTa baseline on the GLUE benchmark without requiring a predefined discrete token vocabulary.

These findings indicate that complex, domain-specific target engineering—such as visual tokenizers or speech unit inventories—is unnecessary for high performance. Unifying the training objective reduces architectural divergence, simplifies multi-modal development pipelines, and lowers maintenance overhead. The results also show that predicting multi-layer, continuous contextual representations makes models exceptionally sample-efficient, especially when labeled training data is scarce.

Organizations developing machine learning systems should consider adopting unified prediction frameworks like data2vec to streamline development across audio, image, and text pipelines. Practitioners should implement multi-layer target averaging rather than relying solely on top-layer predictions. However, teams should note that data2vec still relies on modality-specific input encoders (such as convolutional networks for audio and patch projections for images) and customized masking ratios. Furthermore, when training on highly correlated sequential data like speech, developers must carefully configure learning rate schedules and apply target normalization to avoid representation collapse. Confidence in the framework's core performance is high across standard public benchmarks, but full multi-modal cross-training remains an open area for future operational deployment.

Abstract

While the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to general self-supervised learning, we present data2vec, a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture. Instead of predicting modality-specific targets such as words, visual tokens or units of human speech which are local in nature, data2vec predicts contextualized latent representations that contain information from the entire input. Experiments on the major benchmarks of speech recognition, image classification, and natural language understanding demonstrate a new state of the art or competitive performance to predominant approaches.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Model Architecture
  • 3.2 Masking
  • 3.3 Training Targets
  • 3.4 Objective
  • 4 Experimental Setup
  • 4.1 Computer Vision
  • 4.2 Speech Processing
  • 4.3 Natural Language Processing
  • 5 Results
  • 5.1 Computer Vision
  • 5.2 Speech and Audio Processing
  • 5.3 Natural Language Processing
  • 5.4 Ablations
  • 6 Discussion
  • 7 Conclusion
  • References
  • A Extended speech processing results
  • B Comparison of loss functions
  • C Speech masking parameter ablation

Knowls

  1. Knowl 1 — One masked-prediction self-distillation method across modalities

    model/method

    data2vec uses one self-supervised learning procedure for images, speech, and text. A Transformer teacher encodes the complete input to produce contextualized latent targets; a student with the same model architecture encodes a masked view and is trained to predict those targets at masked positions. The teacher follows the student through an exponential moving average of its parameters. Because the target is a representation of the input rather than a modality-specific label, the learning objective is shared across modalities, while input encoding and masking remain modality-specific.

  2. Knowl 2 — Teacher targets average normalized representations from multiple layers

    model/method

    For an input sequence with LL Transformer blocks, data2vec forms the target at time-step tt by averaging normalized teacher representations from the top KK blocks. Let atla_t^l be the teacher output at time-step tt from block ll, and let a^tl\hat a_t^l be that output after the modality-specific target normalization. The target is

    yt=1K∑l=L−K+1La^tl.y_t = \frac{1}{K}\sum_{l=L-K+1}^{L}\hat a_t^l.

    Here tt indexes a sequence position, ll indexes a Transformer block, and KK is an integer from 11 to LL. The student predicts yty_t at positions masked in its input. Teacher representations are contextualized by self-attention over the full, unmasked input. The usual target feature is the feed-forward network output before the block’s final residual connection.

  3. Knowl 3 — The teacher is an exponentially moving average of student parameters

    equation

    Let θ\theta denote the student parameters and Δ\Delta the teacher parameters. At each update, data2vec updates the teacher as

    Δ←τΔ+(1−τ)θ,\Delta \leftarrow \tau\Delta + (1-\tau)\theta,

    where τ\tau is the teacher’s averaging coefficient. In the scheduled version, τ\tau increases linearly from an initial value τ0\tau_0 to a final value τe\tau_e over the first τn\tau_n training updates and remains at τe\tau_e thereafter. This makes the teacher track the student more quickly early in training and more slowly later. The feature encoder and positional encoder may be shared between teacher and student; the authors found this more efficient and slightly more accurate.

  4. Knowl 4 — Target normalization and regression loss address training stability

    model/method

    data2vec normalizes each teacher layer’s representations before averaging them into targets. For speech, it uses parameter-free instance normalization over the current input sequence; for vision and language, it uses parameter-free layer normalization. The normalization is intended to prevent high-norm layers from dominating the targets and to reduce collapse toward constant representations. The paper reports that collapse is more likely with an excessively large learning rate or too-short warmup, a teacher averaging coefficient that is too low, and highly correlated adjacent targets such as speech representations. Target normalization over the sequence or batch is used to promote variance, particularly for speech.

    For vision and language, the student is generally trained with componentwise Smooth L1 regression. If d=yt−ft(x)d=y_t-f_t(x) is the difference between a target component yty_t and the corresponding student prediction ft(x)f_t(x) for input xx, and β>0\beta>0 is the transition parameter, the loss is

    L(d)={12d2/β,∣d∣≤β,∣d∣−12β,∣d∣>β.\mathcal{L}(d)= \begin{cases} \frac{1}{2}d^2/\beta, & |d|\leq\beta,\\ |d|-\frac{1}{2}\beta, & |d|>\beta. \end{cases}

    The authors used an L2 loss for speech experiments. Smooth L1 is less sensitive to large errors than squared loss, but its transition parameter must be selected.

  5. Knowl 5 — Input encoders and masking are modality-specific

    experimental setup

    The shared data2vec learning procedure uses different input representations and masking strategies for each modality. For vision, a 224×224224\times224 image is divided into 16×1616\times16 patches, linearly embedded, and represented as a sequence of 196 tokens. Block-wise masking covers 60% of patches, with each masked block containing at least 16 adjacent patches; the same resized crop, flip, and color-jittered image is used for teacher and student. For speech, a seven-layer temporal convolutional encoder maps a 16 kHz waveform to 50 Hz representations, with about 20 ms between representations and a 25 ms receptive field. The raw waveform is normalized to zero mean and unit variance. Speech masking samples 6.5% of time-steps as starting positions and masks the following ten time-steps, masking about 49% of a typical sequence. For text, byte-pair encoding produces 50,000 subword types with learned embeddings. The reported language experiments include standard BERT masking and a span-masking variant that masks spans of four tokens with probability 0.35, without leaving tokens unmasked or replacing them with random targets.

  6. Knowl 6 — ImageNet-1K results are strongest among single-model ViT comparisons

    data/table

    The table reports ImageNet-1K validation top-1 accuracy after self-supervised pretraining and supervised fine-tuning. data2vec ViT-B was pretrained for 800 epochs and ViT-L for 1,600 epochs. Results distinguish single-model methods from methods that use additional models, such as a separately trained visual tokenizer or model distillation. Among single models, data2vec scores 84.2 with ViT-B and 86.6 with ViT-L, exceeding the listed single-model comparisons for both sizes. Its ViT-L score is also higher than the listed multiple-model results.

    Method ViT-B ViT-L
    Multiple models
    BEiT 83.2 85.2
    PeCo 84.5 86.5
    Single models
    MoCo v3 83.2 84.1
    DINO 82.8 –
    MAE 83.6 85.9
    SimMIM 83.8 –
    iBOT 83.8 –
    MaskFeat 84.0 85.7
    data2vec 84.2 86.6
  7. Knowl 7 — Speech recognition improves most in low-label settings

    data/table

    The table gives word error rate (WER; lower is better) on the LibriSpeech test-other set after fine-tuning on different amounts of labeled speech. Base systems use 960 hours of unlabeled LibriSpeech audio (LS-960); Large systems use 60,000 hours of Libri-light audio (LL-60K), except WavLM, which uses 94,000 hours (MIX-94K). All listed results use a 4-gram language model for decoding. data2vec Base has lower WER than the listed Base comparisons at every labeled-data amount. For Large systems, data2vec is strongest at 10 minutes and 1 hour, while its 10-hour, 100-hour, and 960-hour results are competitive with the listed models.

    Model Unlabeled data 10m 1h 10h 100h 960h
    Base models; LM: 4-gram
    wav2vec 2.0 LS-960 15.6 11.3 9.5 8.0 6.1
    HuBERT LS-960 15.3 11.3 9.4 8.1 –
    WavLM LS-960 – 10.8 9.2 7.7 –
    data2vec LS-960 12.3 9.1 8.1 6.8 5.5
    Large models; LM: 4-gram
    wav2vec 2.0 LL-60K 10.3 7.1 5.8 4.6 3.6
    HuBERT LL-60K 10.1 6.8 5.5 4.5 3.7
    WavLM MIX-94K – 6.6 5.5 4.6 –
    data2vec LL-60K 8.4 6.3 5.3 4.6 3.7
  8. Knowl 8 — data2vec matches or exceeds the RoBERTa baseline on GLUE

    data/table

    These are GLUE development-set results for individual models fine-tuned separately on each task. MNLI reports matched/unmatched accuracy; MRPC and QQP report the unweighted mean of accuracy and F1; STS-B reports the unweighted mean of Pearson and Spearman correlation; CoLA reports Matthews correlation; the other tasks report accuracy. The average is 82.7 for data2vec, compared with 82.5 for the RoBERTa retraining baseline. Span masking of four BPE tokens raises the reported average to 82.9. The data2vec pretraining setup uses BooksCorpus and English Wikipedia for 1 million updates.

    Model MNLI QNLI RTE MRPC QQP STS-B CoLA SST Avg.
    BERT 84.0/84.4 89.0 61.0 86.3 89.1 89.5 57.3 93.0 80.7
    Baseline (RoBERTa) 84.1/83.9 90.4 69.3 89.0 89.3 88.9 56.8 92.3 82.5
    data2vec 83.2/83.0 90.9 67.0 90.2 89.1 87.2 62.2 91.8 82.7
    data2vec, span masking 82.8/83.4 91.1 69.9 90.0 89.0 87.7 60.3 92.4 82.9
  9. Knowl 9 — Audio event classification improves over listed AudioSet systems

    data/table

    The comparison reports mean average precision (mAP) on the AudioSet evaluation set. Systems are fine-tuned on the 20K subset; the table identifies whether pretraining used AudioSet (AS), Librispeech (LS), or both. data2vec pretrained on AudioSet alone reaches 34.5 mAP, above the listed SSAST, MaskSpec, and MAE-AST results. The pretraining and fine-tuning data differ across comparison systems, so the values are not a fully controlled comparison.

    Method Pretraining data; mAP
    SSAST AS, LS; 31.0
    MaskSpec AS; 32.3
    MAE-AST AS, LS; 30.6
    data2vec AS; 34.5
  10. Knowl 10 — Ablations favor deeper and more contextualized teacher targets

    empirical result

    Ablations across speech, language, and vision show that averaging multiple teacher layers generally improves over using only the top layer (K=1K=1), with a pronounced effect for speech and language and a smaller but still positive effect for vision. Using all layers is generally effective and only slightly worse than a carefully selected layer count. The layer-count experiments used 12-layer Base models; speech was evaluated by WER after 200,000 Librispeech pretraining updates and fine-tuning on 10 hours of labeled data, language by GLUE validation score, and vision by ImageNet top-1 accuracy after 300 pretraining epochs.

    The authors also restricted teacher self-attention to a fraction of the input context when constructing targets, then fine-tuned the resulting model with full context. Increasing the teacher’s accessible context improved downstream performance for both speech and vision, with the best reported performance when the teacher could use the entire input. In a reduced speech setup, the choice of target feature also mattered: using the feed-forward output gave WER 13.1, compared with 100.0 for the self-attention output, 14.8 for the feed-forward output plus residual, and 14.5 for the end-of-block output.

Coverage note — The extended speech dev/test breakdown and the small speech loss- and masking-hyperparameter ablations are omitted because they are supplementary sensitivity results rather than additional central findings.

References

  1. 1.Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text, 2021.
  2. 2.Alayrac, J.-B., Recasens, A., Schneider, R., Arandjelović, R., Ramapuram, J., Fauw, J. D., Smaira, L., Dieleman, S., and Zisserman, A. Self-supervised multimodal versatile networks, 2020.
  3. 3.Ao, J., Wang, R., Zhou, L., Liu, S., Ren, S., Wu, Y., Ko, T., Li, Q., Zhang, Y., Wei, Z., Qian, Y., Li, J., and Wei, F. Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing. arXiv, abs/2110.07205, 2021.
  4. 4.Aytar, Y., Vondrick, C., and Torralba, A. See, hear, and read: Deep aligned representations, 2017.
  5. 5.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv, abs/1607.06450, 2016.
  6. 6.Baade, A., Peng, P., and Harwath, D. Mae-ast: Masked autoencoding audio spectrogram transformer. arXiv, abs/2203.16691, 2022.
  7. 7.Baevski, A., Edunov, S., Liu, Y., Zettlemoyer, L., and Auli, M. Cloze-driven pretraining of self-attention networks. In Proc. of EMNLP, 2019.
  8. 8.Baevski, A., Schneider, S., and Auli, M. vq-wav2vec: Self-supervised learning of discrete speech representations. In Proc. of ICLR, 2020a.
  9. 9.Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proc. of NeurIPS, 2020b.
  10. 10.Baevski, A., Hsu, W.-N., Conneau, A., and Auli, M. Unsupervised speech recognition. In Proc. of NeurIPS, 2021.
  11. 11.Bao, H., Dong, L., and Wei, F. Beit: BERT pre-training of image transformers. arXiv, abs/2106.08254, 2021.
  12. 12.Bardes, A., Ponce, J., and LeCun, Y. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv, abs/2105.04906, 2021.
  13. 13.Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D. The fifth pascal recognizing textual entailment challenge. In Proc. of TAC, 2009.
  14. 14.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Proc. of NeurIPS, 2020.
  15. 15.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. arXiv, abs/2006.09882, 2020.
  16. 16.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. arXiv, abs/2104.14294, 2021.
  17. 17.Cer, D. M., Diab, M. T., Agirre, E., Lopez-Gazpio, I., and Specia, L. Semeval-2017 task 1: Semantic textual similarity - multilingual and cross-lingual focused evaluation. In Proc. of SemEval, 2018.
  18. 18.Chang, H.-J., wen Yang, S., and yi Lee, H. Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert. arXiv, abs/2110.01900, 2021.
  19. 19.Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y., Qian, Y., Wu, J., Zeng, M., and Wei, F. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. arXiv, abs/2110.13900, 2021a.
  20. 20.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. arXiv, abs/2002.05709, 2020.
  21. 21.Chen, X., Xie, S., and He, K. An empirical study of training self-supervised vision transformers. arXiv, abs/2104.02057, 2021b.
  22. 22.Chong, D., Wang, H., Zhou, P., and Zeng, Q. Masked spectrogram prediction for self-supervised audio pre-training. arXiv, abs/2204.12768, 2022.
  23. 23.Chung, Y., Hsu, W., Tang, H., and Glass, J. R. An unsupervised autoregressive model for speech representation learning. Proc. of Interspeech, 2019.
  24. 24.Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y. W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. arXiv, abs/2108.06209, 2021.
  25. 25.Dagan, I., Glickman, O., and Magnini, B. The pascal recognizing textual entailment challenge. Machine learning challenges, evaluating predictive uncertainty, visual object classification, and recognizing textual entailment, pp. 177–190, 2006.
  26. 26.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In Proc. of CVPR, 2009.
  27. 27.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. Proc. of NAACL, 2019.
  28. 28.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In Proc. of IWP, 2005.
  29. 29.Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N. Peco: Perceptual codebook for bert pre-training of vision transformers, 2022.
  30. 30.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv, abs/2010.11929, 2020.
  31. 31.Eloff, R., Nortje, A., van Niekerk, B., Govender, A., Nortje, L., Pretorius, A., Van Biljon, E., van der Westhuizen, E., van Staden, L., and Kamper, H. Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks. arXiv, abs/1904.07556, 2019.
  32. 32.Friston, K. The free-energy principle: a unified brain theory? Nature reviews neuroscience, 2010.
  33. 33.Friston, K. and Kiebel, S. Predictive coding under the free-energy principle. Philosophical transactions of the Royal Society: Biological sciences, 2009.
  34. 34.Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. Audio set: An ontology and human-labeled dataset for audio events. In Proc. of ICASSP, 2017.
  35. 35.Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. The pascal recognizing textual entailment challenge. Proc. of the ACL-PASCAL workshop on textual entailment and paraphrasing, 2007.
  36. 36.Gong, Y., Lai, C.-I. J., Chung, Y.-A., and Glass, J. Ssast: Self-supervised audio spectrogram transformer. arXiv, abs/2110.09784, 2021.
  37. 37.Grill, J.-B., Strub, F., Altche, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. Bootstrap your own latent: A new approach to self-supervised learning. arXiv, abs/2006.07733, 2020.
  38. 38.Haim, R. B., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I. The pascal recognising textual entailment challenge. Lecture Notes in Computer Science, 2006.
  39. 39.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. arXiv, abs/1911.05722, 2019.
  40. 40.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. arXiv, abs/2111.06377, 2021.
  41. 41.Hsu, W.-N., Tsai, Y.-H. H., Bolte, B., Salakhutdinov, R., and Mohamed, A. Hubert: How much can a bad teacher benefit ASR pre-training? In Proc. of ICASSP, 2021.
  42. 42.Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Deep networks with stochastic depth. arXiv, abs/1603.09382, 2016.
  43. 43.Jaegle, A., Borgeaud, S., Alayrac, J.-B., Doersch, C., Ionescu, C., Ding, D., Koppula, S., Zoran, D., Brock, A., Shelhamer, E., Henaff, O., Botvinick, M. M., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver io: A general architecture for structured inputs & outputs. arXiv, abs/2107.14795, 2021a.
  44. 44.Jaegle, A., Gimeno, F., Brock, A., Zisserman, A., Vinyals, O., and Carreira, J. Perceiver: General perception with iterative attention. arXiv, abs/2103.03206, 2021b.
  45. 45.Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. Tinybert: Distilling bert for natural language understanding. arXiv, abs/1909.10351, 2020.
  46. 46.Jing, L., Vincent, P., LeCun, Y., and Tian, Y. Understanding dimensional collapse in contrastive self-supervised learning, 2021.
  47. 47.Kahn, J., Riviere, M., Zheng, W., Kharitonov, E., Xu, Q., Mazare, P., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E. Libri-light: A benchmark for asr with limited or no supervision. arXiv, abs/1912.07875, 2019.
  48. 48.Kahn, J. et al. Libri-light: A benchmark for asr with limited or no supervision. In Proc. of ICASSP, 2020.
  49. 49.Kingma, D. P. and Ba, J. Adam: A Method for Stochastic Optimization. In Proc. of ICLR, 2015.
  50. 50.Lample, G., Denoyer, L., and Ranzato, M. Unsupervised machine translation using monolingual corpora only. In Proc. of ICLR, 2018.
  51. 51.Likhomanenko, T., Xu, Q., Kahn, J., Synnaeve, G., and Collobert, R. slimipl: Language-model-free iterative pseudo-labeling. arXiv, abs/2010.11524, 2021.
  52. 52.Ling, S., Liu, Y., Salazar, J., and Kirchhoff, K. Deep contextualized acoustic representations for semi-supervised speech recognition. In Proc. of ICASSP, 2020.
  53. 53.Liu, A. T., Li, S.-W., and Lee, H.-y. Tera: Self-supervised learning of transformer encoder representation for speech. IEEE/ACM Trans. on Audio, Speech, and Language Processing, 2021.
  54. 54.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  55. 55.Loshchilov, I. and Hutter, F. SGDR: stochastic gradient descent with restarts. arXiv, abs/1608.03983, 2016.
  56. 56.Manohar, V., Likhomanenko, T., Xu, Q., Hsu, W.-N., Collobert, R., Saraf, Y., Zweig, G., and Mohamed, A. Kaizen: Continuously improving teacher using exponential moving average for semi-supervised speech recognition, 2021.
  57. 57.McCann, B., Bradbury, J., Xiong, C., and Socher, R. Learned in translation: Contextualized word vectors. arXiv, abs/1708.00107, 2017.
  58. 58.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In Proc. of NAACL System Demonstrations, 2019.
  59. 59.Park, D. S., Zhang, Y., Jia, Y., Han, W., Chiu, C.-C., Li, B., Wu, Y., and Le, Q. V. Improved noisy student training for automatic speech recognition. Proc. of Interspeech, 2020.
  60. 60.Pasad, A., Chou, J.-C., and Livescu, K. Layer-wise analysis of a self-supervised speech representation model. arXiv, abs/2107.04734, 2021.
  61. 61.Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. Deep contextualized word representations. In Proc. of ACL, 2018.
  62. 62.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf, 2018.
  63. 63.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv, abs/2103.00020, 2021a.
  64. 64.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv, abs/2103.00020, 2021b.
  65. 65.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100, 000+ questions for machine comprehension of text. arXiv, abs/1606.05250, 2016.
  66. 66.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. arXiv, abs/2102.12092, 2021.
  67. 67.Schneider, S., Baevski, A., Collobert, R., and Auli, M. wav2vec: Unsupervised pre-training for speech recognition. In Proc. of Interspeech, 2019.
  68. 68.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proc. of ACL, 2016.
  69. 69.Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. FLAVA: A foundational language and vision alignment model. arXiv, abs/2112.04482, 2021.
  70. 70.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proc. of EMNLP, 2013.
  71. 71.Srivastava, S., Wang, Y., Tjandra, A., Kumar, A., Liu, C., Singh, K., and Saraf, Y. Conformer-based self-supervised learning for non-speech audio tasks. arXiv, abs/2110.07313, 2021.
  72. 72.Tarvainen, A. and Valpola, H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, 2018.
  73. 73.Tay, Y., Tran, V. Q., Ruder, S., Gupta, J., Chung, H. W., Bahri, D., Qin, Z., Baumgartner, S., Yu, C., and Metzler, D. Charformer: Fast character transformers via gradient-based subword tokenization. arXiv, abs/2106.12672, 2021.
  74. 74.Tokozume, Y., Ushiku, Y., and Harada, T. Learning from between-class examples for deep sound recognition. arXiv, abs/1711.10282, 2017.
  75. 75.Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S. M. A., Vinyals, O., and Hill, F. Multimodal few-shot learning with frozen language models, 2021.
  76. 76.Ulyanov, D., Vedaldi, A., and Lempitsky, V. S. Instance normalization: The missing ingredient for fast stylization. arXiv, abs/1607.08022, 2016.
  77. 77.van den Oord, A., Vinyals, O., et al. Neural discrete representation learning. In Proc. of NeurIPS, 2017.
  78. 78.van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. Proc. of NIPS, 2018.
  79. 79.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Proc. of NIPS, 2017.
  80. 80.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv, abs/1804.07461, 2018a.
  81. 81.Wang, H., Ma, S., Dong, L., Huang, S., Zhang, D., and Wei, F. Deepnet: Scaling transformers to 1,000 layers. arXiv, abs/2203.00555, 2022.
  82. 82.Wang, W., Bao, H., Dong, L., and Wei, F. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv, abs/2111.02358, 2021.
  83. 83.Wang, Y., Li, J., and Metze, F. A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling. arXiv, abs/1810.09050, 2018b.
  84. 84.Warstadt, A., Singh, A., and Bowman, S. Corpus of linguistic acceptability. https://nyu-mll.github.io/CoLA, 2018.
  85. 85.Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. arXiv, abs/2112.09133, 2021.
  86. 86.Williams, A., Nangia, N., and Bowman, S. R. A broad-coverage challenge corpus for sentence understanding through inference. In Proc. of NAACL, 2018.
  87. 87.Wu, Z., Wang, S., Gu, J., Khabsa, M., Sun, F., and Ma, H. CLEAR: contrastive learning for sentence representation. arXiv, abs/2012.15466, 2020.
  88. 88.Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. Simmim: A simple framework for masked image modeling. arXiv, abs/2111.09886, 2021.
  89. 89.Xu, Q., Likhomanenko, T., Kahn, J., Hannun, A., Synnaeve, G., and Collobert, R. Iterative pseudo-labeling for speech recognition. Proc. of Interspeech, 2020.
  90. 90.Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C. Byt5: Towards a token-free future with pre-trained byte-to-byte models. arXiv, abs/2105.13626, 2021.
  91. 91.Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V. Xlnet: Generalized autoregressive pretraining for language understanding. arXiv, abs/1906.08237, 2019.
  92. 92.Zhang, Y., Qin, J., Park, D. S., Han, W., Chiu, C.-C., Pang, R., Le, Q. V., and Wu, Y. Pushing the limits of semi-supervised learning for automatic speech recognition. Proc. of NeurIPS SAS Workshop, 2020.
  93. 93.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. ibot: Image bert pre-training with online tokenizer, 2021.
  94. 94.Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. arXiv, abs/1506.06724, 2015.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/