BEATs: Audio Pre-Training with Acoustic Tokenizers

Sanyuan ChenYu WuChengyi WangShujie LiuDaniel TompkinsZhuo ChenWanxiang CheXiangzhan YuFuru Wei

article2023ICML549 citations

Proposes an iterative self-supervised pre-training framework that jointly optimizes an acoustic tokenizer and an audio transformer via discrete label prediction, achieving state-of-the-art classification performance on AudioSet-2M and ESC-50 while using fewer parameters and less data than previous reconstruction-based methods.

Listen

Building powerful artificial intelligence models to understand general audio is challenging due to the diverse, unstructured nature of environmental sounds, background noise, and speech. While existing self-supervised audio models typically rely on reconstructing low-level acoustic details—such as raw spectrograms—this approach often retains background noise and fails to extract high-level semantic meaning. The article addresses this limitation by developing an iterative audio pre-training framework called BEATS (Bidirectional Encoder representation from Audio Transformers), which shifts the training objective from reconstructing low-level acoustic details to predicting high-level discrete acoustic tokens.

The framework pairs an acoustic tokenizer with a transformer-based audio encoder in an alternating, iterative training loop. Initially, a random-projection tokenizer creates baseline discrete labels. The audio encoder learns to predict these discrete tokens for masked audio segments. Once trained, the encoder acts as a teacher to guide and refine a new, self-distilled acoustic tokenizer through knowledge distillation. The newly refined tokenizer generates richer semantic labels to train the subsequent audio encoder, and the process repeats. The article pre-trained the framework on roughly two million clips from the AudioSet database and evaluated the resulting models across six standard audio and speech classification benchmarks.

The experimental findings show that BEATS consistently outperforms previous approaches while using fewer resources. First, single BEATS models established new state-of-the-art results on major benchmarks, achieving a 48.6% mean Average Precision (mAP) on the full AudioSet-2M dataset and a 25% relative error reduction on the ESC-50 environmental sound benchmark (reaching 98.1% accuracy). Second, BEATS outperformed previous leading models that required over three times as many parameters (90 million versus 304 million) or extensive pre-training on external visual datasets like ImageNet. Third, ensembling ten fine-tuned BEATS models achieved a record 50.6% mAP on AudioSet-2M without using any external data. Finally, visualization and disturbance analyses confirmed that the acoustic tokenizers successfully ignore background noise and reverberation, isolating meaningful semantic content far better than reconstruction-based methods.

These findings indicate that discrete label prediction enables more semantically focused and computationally efficient audio representation learning. By matching or exceeding larger systems with significantly smaller model architectures, this approach lowers deployment costs and reduces dependence on massive supervised labels or cross-modal pre-training data. Furthermore, using discrete tokens aligns audio pre-training methods with those used in natural language, computer vision, and speech processing, offering a viable pathway toward building unified multimodal foundation models.

Stakeholders and engineering teams should consider adopting the released BEATS pre-trained models and acoustic tokenizers to optimize sound classification workflows, lower labeling overhead, and accelerate downstream application development. For future work, development efforts should focus on scaling the architecture beyond 90 million parameters, expanding pre-training data beyond AudioSet to broader audio domains, and integrating the tokenizer framework into multimodal systems combining audio, text, and vision. While the reported performance gains are statistically robust across multiple standard benchmarks, the main practical trade-off is the linear increase in training compute required during the iterative tokenizer-model optimization cycles.

Cover for BEATs: Audio Pre-Training with Acoustic Tokenizers

Abstract

We introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. BEATS
  • 3.1. Iterative Audio Pre-training
  • 3.2. Acoustic Tokenizers
  • 3.2.1. COLD START: RANDOM-PROJECTION TOKENIZER
  • 3.2.2. ITERATION: SELF-DISTILLED TOKENIZER
  • 3.3. Audio SSL Model
  • 3.3.1. BACKBONE
  • 3.3.2. PRE-TRAINING
  • 3.3.3. FINE-TUNING
  • 4. Experiment
  • 4.1. Setup
  • 4.2. Comparing with the SOTA Single Models
  • 4.3. Comparing Different BEATS Tokenizers
  • 4.4. Comparing Different Pre-Training Targets via Visualization
  • 4.5. Comparing with the SOTA Ensemble Models
  • 5. Conclusion, Limitations, and Future Work
  • References
  • A. Convergence Analysis
  • A.1. Mathematical Formulation
  • A.2. Convergence Proof
  • B. Datasets
  • C. Hyperparamter Settings

Knowls

  1. Knowl 1 — Iterative Audio Pre-Training Framework (BEATS)

    model/method

    The BEATS (Bidirectional Encoder representation from Audio Transformers) framework optimizes an acoustic tokenizer and an audio self-supervised learning (SSL) model alternately across multiple iterations:

    1. Cold-Start Iteration (t=1t=1): A frozen random-projection tokenizer clusters continuous audio patch features into discrete acoustic labels. An audio SSL model is pre-trained to predict these discrete labels on masked audio patches.

    2. Subsequent Iterations (t≥2t \ge 2): A self-distilled acoustic tokenizer is trained via knowledge distillation using the audio SSL model from iteration t−1t-1 (which can be either a pre-trained SSL model or a model fine-tuned on downstream labeled data) as the teacher. The newly trained tokenizer then quantizes the unlabeled audio dataset to produce refined discrete semantic labels. A new audio SSL model is trained from scratch using the updated discrete labels as masked prediction targets.

    This alternating optimization process allows the acoustic tokenizer to learn high-level semantic abstractions from the SSL teacher while enabling the SSL model to learn rich discrete representations from the improved tokenizer.

  2. Knowl 2 — Self-Distilled Acoustic Tokenizer Architecture and Training Loss

    model/method

    From the second pre-training iteration of BEATS, a self-distilled acoustic tokenizer converts continuous audio features into semantic discrete tokens using knowledge distillation from an audio SSL teacher model obtained in the previous iteration.

    Given an input audio clip represented as a sequence of regular 2D spectrogram patches X={xt}t=1TX = \{x_t\}_{t=1}^T:

    1. A 12-layer Transformer encoder maps XX to continuous latent representations E={et}t=1TE = \{e_t\}_{t=1}^T.

    2. Each vector ete_t is quantized by finding the nearest neighbor in a learnable codebook V={vi}i=1KV = \{v_i\}_{i=1}^K containing K=1024K=1024 embeddings of dimension 256 using ℓ2\ell_2-normalized Euclidean distance: z^t=arg⁡min⁡i∈{1,…,K}∥ℓ2(vi)−ℓ2(et)∥22\hat{z}_t = \arg\min_{i \in \{1,\dots,K\}} \|\ell_2(v_i) - \ell_2(e_t)\|_2^2 The quantized sequence is Eq={vz^t}t=1TE^q = \{v_{\hat{z}_t}\}_{t=1}^T. During backpropagation, straight-through gradient estimation copies gradients directly from EqE^q to EE, and the codebook embeddings are updated via exponential moving average (EMA).

    3. A 3-layer Transformer estimator takes EqE^q as input and predicts representation vectors {ot}t=1T\{o_t\}_{t=1}^T matching the teacher model's final-layer representations {o^t}t=1T\{\hat{o}_t\}_{t=1}^T.

    The training objective maximizes cosine similarity to the teacher outputs while minimizing vector quantization error with stop-gradient (sg[⋅]\text{sg}[\cdot]): Ltokenizer=∑X∈D∑t=1Tcos⁡(ot,o^t)−∥sg[ℓ2(et)]−ℓ2(vz^t)∥22−∥ℓ2(et)−sg[ℓ2(vz^t)]∥22\mathcal{L}_{\text{tokenizer}} = \sum_{X \in \mathcal{D}} \sum_{t=1}^T \cos(o_t, \hat{o}_t) - \|\text{sg}[\ell_2(e_t)] - \ell_2(v_{\hat{z}_t})\|_2^2 - \|\ell_2(e_t) - \text{sg}[\ell_2(v_{\hat{z}_t})]\|_2^2 where D\mathcal{D} is the pre-training dataset and cos⁡(⋅,⋅)\cos(\cdot, \cdot) is cosine similarity. During inference, the estimator is discarded and the encoder-codebook pair is used to generate patch-level discrete labels Z^={z^t}t=1T\hat{Z} = \{\hat{z}_t\}_{t=1}^T.

  3. Knowl 3 — Masked Audio Modeling Pre-Training for BEATS

    model/method

    The BEATS audio self-supervised learning (SSL) model uses a 12-layer Vision Transformer (ViT) backbone (768 hidden dimension, 8 attention heads, 90M parameters) equipped with a convolution-based relative position embedding layer, gated relative position bias, and DeepNorm initialization.

    The input waveform is sampled at 16 kHz and converted to 128-dimensional Mel-filter bank features (25 ms Povey window, 10 ms hop size, normalized to zero mean and standard deviation 0.5), which are partitioned into non-overlapping 16×1616 \times 16 time-frequency patches X={xt}t=1TX = \{x_t\}_{t=1}^T.

    During Masked Audio Modeling (MAM) pre-training:

    1. A random subset M⊂{1,…,T}\mathcal{M} \subset \{1, \dots, T\} comprising 75% of the patch sequence is masked.

    2. Only the unmasked patches XU={xt:t∉M}X^U = \{x_t : t \notin \mathcal{M}\} are linearly projected and fed into the ViT encoder to produce unmasked representations RU={rt:t∉M}R^U = \{r_t : t \notin \mathcal{M}\}, reducing computational cost.

    3. The unmasked representations are recombined with zero embeddings at the masked positions: {rt:t∉M}∪{0:t∈M}\{r_t : t \notin \mathcal{M}\} \cup \{0 : t \in \mathcal{M}\}.

    4. A linear label predictor predicts discrete target labels Z^={z^t}t=1T\hat{Z} = \{\hat{z}_t\}_{t=1}^T generated by the acoustic tokenizer.

    5. The model is optimized using cross-entropy loss over the masked positions: LMAM=−∑t∈Mlog⁡p(z^t∣XU)\mathcal{L}_{\text{MAM}} = -\sum_{t \in \mathcal{M}} \log p(\hat{z}_t | X^U)

  4. Knowl 4 — Random-Projection Acoustic Tokenizer for Cold Start

    model/method

    In the initial pre-training iteration of BEATS (t=1t=1), when no trained teacher model is available, a random-projection acoustic tokenizer is used to discretize continuous audio spectrogram patches.

    Given an input patch sequence X={xt}t=1TX = \{x_t\}_{t=1}^T, the tokenizer uses a randomly initialized, frozen linear projection matrix WW and a frozen codebook of K=1024K=1024 randomly initialized vectors V={vi}i=1K⊂R256V = \{v_i\}_{i=1}^K \subset \mathbb{R}^{256}. The patch xtx_t is linearly projected to WxtW x_t, and the discrete token z^t\hat{z}_t is assigned as the index of the nearest codebook vector: z^t=arg⁡min⁡i∈{1,…,K}∥vi−Wxt∥22\hat{z}_t = \arg\min_{i \in \{1,\dots,K\}} \|v_i - W x_t\|_2^2 This non-learned clustering step provides the cold-start target labels needed to train the first-iteration audio SSL model.

  5. Knowl 5 — Expectation-Maximization Formulation and Convergence of Iterative Pre-training

    theoretical result

    The alternating pre-training of the acoustic tokenizer parameters δ\delta and the audio SSL model parameters θ\theta can be formulated as an Expectation-Maximization (EM) procedure maximizing the observable data log-likelihood log⁡p(X∣θ,δ)\log p(X | \theta, \delta) with latent variables given by the discrete labels ZZ and SSL hidden representations RR.

    Under this formulation:

    1. Tokenizer optimization (E-step over RR, M-step over δ\delta): Training the tokenizer at iteration t+1t+1 to maximize the joint distribution p(X,R∣θ(t),δ)p(X, R | \theta^{(t)}, \delta) with R∼p(R∣X,θ(t))R \sim p(R | X, \theta^{(t)}) guarantees that the likelihood is non-decreasing: log⁡p(X∣θ(t),δ(t+1))≥log⁡p(X∣θ(t),δ(t))\log p(X | \theta^{(t)}, \delta^{(t+1)}) \ge \log p(X | \theta^{(t)}, \delta^{(t)})

    2. SSL model optimization (E-step over ZZ, M-step over θ\theta): Training the SSL model at iteration t+1t+1 to maximize p(X,Z∣θ,δ(t+1))p(X, Z | \theta, \delta^{(t+1)}) with Z∼p(Z∣X,δ(t+1))Z \sim p(Z | X, \delta^{(t+1)}) guarantees: log⁡p(X∣θ(t+1),δ(t+1))≥log⁡p(X∣θ(t),δ(t+1))\log p(X | \theta^{(t+1)}, \delta^{(t+1)}) \ge \log p(X | \theta^{(t)}, \delta^{(t+1)})

    Combining both steps yields the non-decreasing log-likelihood property across iterations: log⁡p(X∣θ(t+1),δ(t+1))≥log⁡p(X∣θ(t),δ(t))\log p(X | \theta^{(t+1)}, \delta^{(t+1)}) \ge \log p(X | \theta^{(t)}, \delta^{(t)}) This ensures convergence of the iterative pre-training framework to a local maximum or saddle point of the data log-likelihood.

  6. Knowl 6 — Downstream Fine-Tuning and Evaluation Methodology

    experimental setup

    For downstream task adaptation, the pre-trained BEATS label predictor is discarded, and a linear classification head with mean pooling is placed on top of the ViT encoder: p(C)=Softmax(MeanPool(WcR))p(C) = \text{Softmax}(\text{MeanPool}(W_c R)) where R={rt}t=1TR = \{r_t\}_{t=1}^T are the encoded representations of the full patch sequence XX (augmented with SpecAugment), and WcW_c is a task-specific linear projection matrix. Cross-entropy loss is used for single-label tasks, and binary cross-entropy loss is used for multi-label tasks or mixup augmentation.

    Evaluation tasks include:

    • AudioSet-2M (AS-2M): 2M 10-second YouTube clips, 527 classes (multi-label), evaluated using mean Average Precision (mAP).
    • AudioSet-20K (AS-20K): 21K balanced subset of AudioSet, evaluated using mAP.
    • ESC-50: 2,000 5-second environmental recordings across 50 classes, evaluated using 5-fold cross-validation accuracy.
    • Speech Commands V1 (KS1): 12 classes (10 keyword classes, 1 silence class, 1 unknown class), evaluated by classification accuracy under the SUPERB benchmark split.
    • Speech Commands V2 (KS2): 105,829 1-second clips across 35 word classes, evaluated by classification accuracy.
    • IEMOCAP (ER): Emotion recognition across 4 emotion classes (~12 hours), evaluated by 5-fold cross-validation accuracy.
  7. Knowl 7 — Audio and Speech Classification Benchmark Performance

    data/table

    Single-model classification results comparing BEATS models across pre-training iterations with prior supervised and self-supervised models on six audio and speech benchmarks. Evaluation metrics are mean Average Precision (mAP) for AS-2M and AS-20K, and accuracy (%) for ESC-50, KS1, KS2, and ER.

    Model # Param Pre-train Data AS-2M AS-20K ESC-50 KS1 KS2 ER
    Out-of-Domain / In-Domain Supervised Pre-Training
    AST 86M ImageNet 45.9 34.7 88.7 95.5 98.1 56.0
    PaSST 86M ImageNet+AudioSet 47.1 - 96.8 - - -
    HTS-AT 31M ImageNet+AudioSet 47.1 - 97.0 - 98.0 -
    Audio-MAE (Sup.) 86M AudioSet - - 97.4 - - -
    Self-Supervised Pre-Training
    SS-AST 89M AudioSet+LibriSpeech - 31.0 88.8 96.0 98.0 59.6
    MSM-MAE 86M AudioSet - - 85.6 - 87.3 -
    MaskSpec 86M AudioSet 47.1 32.3 89.6 - 97.7 -
    MAE-AST 86M AudioSet+LibriSpeech - 30.6 90.0 95.8 97.9 59.8
    data2vec 94M AudioSet - 34.5 - - - -
    Audio-MAE 86M AudioSet 47.3 37.1 94.1 96.9 98.3 -
    Audio-MAE Large 304M AudioSet 47.4 37.6 - - - -
    CAV-MAE 86M AudioSet+ImageNet 44.9 34.2 - - - -
    BEATS (Ours)
    BEATSiter1\text{BEATS}_{\text{iter1}} 90M AudioSet 47.9 36.0 94.0 98.0 98.3 65.9
    BEATSiter2\text{BEATS}_{\text{iter2}} 90M AudioSet 48.1 38.3 95.1 97.7 98.3 66.1
    BEATSiter3\text{BEATS}_{\text{iter3}} 90M AudioSet 48.0 38.3 95.6 97.7 98.3 64.5
    BEATSiter3+\text{BEATS}_{\text{iter3+}} 90M AudioSet 48.6 38.9 98.1 98.1 98.1 65.0

    BEATSiter1\text{BEATS}_{\text{iter1}} (trained with the random-projection tokenizer) outperforms prior SSL methods on five tasks. Self-distilled iterations improve performance further, and BEATSiter3+\text{BEATS}_{\text{iter3+}} (which uses a supervised fine-tuned teacher) achieves a state-of-the-art single-model mAP of 48.6% on AS-2M (outperforming 304M-parameter Audio-MAE Large at 47.4%) and 98.1% accuracy on ESC-50 (with AS-2M pre-training), reducing relative classification error by 25%.

  8. Knowl 8 — AudioSet-2M State-of-the-Art Ensemble Performance

    empirical result

    Model ensembling of fine-tuned BEATS models on AudioSet-2M (AS-2M) demonstrates state-of-the-art performance without utilizing external supervised pre-training data such as ImageNet:

    • BEATS (5 models): An ensemble of five fine-tuned BEATS models achieves 50.4% mAP on AS-2M, outperforming prior best ensemble models (PaSST with ImageNet pre-training at 49.6% mAP, HTS-AT at 48.7% mAP, AST at 48.5% mAP, and PSLA at 47.4% mAP).
    • BEATS (10 models): Re-running AS-2M fine-tuning of the five BEATS SSL models with a learning rate of 5×10−55 \times 10^{-5} for 100k steps and ensembling all ten models yields a new state-of-the-art 50.6% mAP on AS-2M.
  9. Knowl 9 — Ablation of Tokenizer Architectures and Teacher Supervision

    data/table

    Detailed comparison of different BEATS tokenizer variants and teacher supervision sources across audio classification benchmarks. ESC-50 results are reported without additional supervised pre-training on AudioSet-2M.

    Model Tokenizer Type Tokenizer Teacher SSL Data SL Data AS-2M (mAP) AS-20K (mAP) ESC-50 (Acc)
    BEATSiter1\text{BEATS}_{\text{iter1}} Random-Projection N/A AS - 47.9 36.0 94.0
    BEATSiter2\text{BEATS}_{\text{iter2}} Self-Distilled BEATSiter1\text{BEATS}_{\text{iter1}} AS - 48.1 38.3 95.1
    BEATSiter3\text{BEATS}_{\text{iter3}} Self-Distilled BEATSiter2\text{BEATS}_{\text{iter2}} AS - 48.0 38.3 95.6
    BEATSiter3+\text{BEATS}_{\text{iter3+}} Self-Distilled BEATSiter2\text{BEATS}_{\text{iter2}} (fine-tuned on AS-20K) AS AS-20K 48.0 38.9 96.2
    BEATSiter3+\text{BEATS}_{\text{iter3+}} Self-Distilled BEATSiter2\text{BEATS}_{\text{iter2}} (fine-tuned on AS-2M) AS AS 48.6 41.8* 97.1

    *Note: The 41.8% AS-20K score uses AS-2M supervised data during tokenizer pre-training and AS-20K for downstream fine-tuning.

    Key takeaways:

    1. Self-distilled tokenizers provide substantial improvements over the random-projection tokenizer, especially on data-scarce setups (AS-20K mAP increases from 36.0% to 38.3%, ESC-50 accuracy increases from 94.0% to 95.1%).
    2. The self-distilled tokenizer performance is robust across iterations of self-supervised teachers (BEATSiter1\text{BEATS}_{\text{iter1}} vs. BEATSiter2\text{BEATS}_{\text{iter2}} yield matching 38.3% mAP on AS-20K).
    3. Using a supervised fine-tuned teacher model distills task-specific semantic information into the tokenizer, achieving the highest performance across all tasks (48.6% mAP on AS-2M, 97.1% accuracy on ESC-50 without AS-2M supervised pre-training).
  10. Knowl 10 — Semantic Clustering and Perturbation Robustness of BEATS Token Targets

    empirical result

    T-SNE 2D visualizations comparing pre-training targets on ESC-50 samples subjected to Room Impulse Response (RIR) reverberation and Deep Noise Suppression (DNS) background noise show distinct differences between reconstruction-based targets and BEATS discrete targets:

    • Reconstruction-based SSL targets (continuous acoustic Mel-filter bank features): Highly sensitive to random acoustic variations. Disturbed variants of the same audio sample scatter widely across feature space, while distinct semantic classes overlap closely. This demonstrates that continuous reconstruction objectives focus heavily on low-level time-frequency details and lack high-level semantic abstraction.

    • BEATSiter3\text{BEATS}_{\text{iter3}} discrete targets (self-supervised teacher distillation): Cluster perturbed versions of audio samples belonging to the same semantic class together, filtering out background noise and reverberation.

    • BEATSiter3+\text{BEATS}_{\text{iter3+}} discrete targets (supervised fine-tuned teacher distillation): Form tight, well-separated clusters corresponding strictly to semantic class identities, ignoring acoustic perturbations entirely and providing robust semantic targets for masked audio modeling.

  11. Knowl 11 — Computational, Data, and Scale Limitations of BEATS

    limitation

    The BEATS framework has three primary stated limitations:

    1. Computational Overhead: The iterative pre-training framework introduces a computational overhead that scales linearly with the number of iteration rounds (pre-training each 90M SSL model requires ~75 hours on 16 Tesla V100 GPUs, and training each self-distilled tokenizer requires ~45 hours on 8 Tesla V100 GPUs).

    2. Data Coverage: Pre-training relies solely on AudioSet-2M, which limits data diversity and coverage across broader acoustic, music, and multilingual speech domains.

    3. Model Scalability: Model size in this study is constrained to 90M parameters (96M total including heads), which is considerably smaller than modern foundation models in speech, vision, and natural language processing.

Coverage note — None was omitted; all key architectural components, training algorithms, theoretical convergence results, empirical benchmark performances, tokenizer ablation comparisons, target visualizations, and stated limitations are covered.

References

  1. 1.Al-Tahan, H. and Mohsenzadeh, Y. Clar: Contrastive learning of auditory representations. In International Conference on Artificial Intelligence and Statistics, pp. 2530–2538. PMLR, 2021.
  2. 2.Baade, A., Peng, P., and Harwath, D. Mae-ast: Masked autoencoding audio spectrogram transformer. arXiv preprint arXiv:2203.16691, 2022.
  3. 3.Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33:12449–12460, 2020.
  4. 4.Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022.
  5. 5.Bao, H., Dong, L., Piao, S., and Wei, F. Beit: Bert pretraining of image transformers. In International Conference on Learning Representations, 2021.
  6. 6.Bao, H., Wang, W., Dong, L., and Wei, F. Vl-beit: Generative vision-language pretraining. arXiv preprint arXiv:2206.01127, 2022.
  7. 7.Busso, C., Bulut, M., Lee, C.-C., Kazemzadeh, A., Mower, E., Kim, S., Chang, J. N., Lee, S., and Narayanan, S. S. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42(4):335–359, 2008.
  8. 8.Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., and Dubnov, S. Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 646–650. IEEE, 2022a.
  9. 9.Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022b.
  10. 10.Chi, Z., Huang, S., Dong, L., Ma, S., Zheng, B., Singhal, S., Bajaj, P., Song, X., Mao, X.-L., Huang, H.-Y., et al. Xlm-e: Cross-lingual language model pre-training via electra. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6170–6182, 2022.
  11. 11.Chiu, C.-C., Qin, J., Zhang, Y., Yu, J., and Wu, Y. Self-supervised learning with random-projection quantizer for speech recognition. arXiv preprint arXiv:2202.01855, 2022.
  12. 12.Chong, D., Wang, H., Zhou, P., and Zeng, Q. Masked spectrogram prediction for self-supervised audio pre-training. arXiv preprint arXiv:2204.12768, 2022.
  13. 13.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  14. 14.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, 2019.
  15. 15.Dieleman, S., van den Oord, A., and Simonyan, K. The challenge of realistic music generation: modelling raw audio at scale. Advances in Neural Information Processing Systems, 31, 2018.
  16. 16.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  17. 17.Elizalde, B., Deshmukh, S., Ismail, M. A., and Wang, H. Clap: Learning audio concepts from natural language supervision. arXiv preprint arXiv:2206.04769, 2022.
  18. 18.Fonseca, E., Ortego, D., McGuinness, K., O’Connor, N. E., and Serra, X. Unsupervised contrastive learning of sound event representations. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 371–375. IEEE, 2021.
  19. 19.Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 776–780. IEEE, 2017.
  20. 20.Gong, Y., Chung, Y.-A., and Glass, J. Ast: Audio spectrogram transformer. arXiv preprint arXiv:2104.01778, 2021a.
  21. 21.Gong, Y., Chung, Y.-A., and Glass, J. Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3292–3306, 2021b.
  22. 22.Gong, Y., Lai, C.-I., Chung, Y.-A., and Glass, J. Ssast: Self-supervised audio spectrogram transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 10699–10709, 2022a.
  23. 23.Gong, Y., Rouditchenko, A., Liu, A. H., Harwath, D., Karlinsky, L., Kuehne, H., and Glass, J. Contrastive audio-visual masked autoencoder. arXiv preprint arXiv:2210.07839, 2022b.
  24. 24.Guzhov, A., Raue, F., Hees, J., and Dengel, A. Audioclip: Extending clip to image, text and audio. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 976–980. IEEE, 2022.
  25. 25.Harb, H. and Chen, L. A general audio classifier based on human perception motivated model. Multimedia Tools and Applications, 34:375–395, 2007.
  26. 26.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000–16009, 2022.
  27. 27.Hinton, G., Vinyals, O., Dean, J., et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  28. 28.Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  29. 29.Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. Panns: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2880–2894, 2020.
  30. 30.Koutini, K., Schluter, J., Eghbal-zadeh, H., and Widmer, G. Efficient training of audio transformers with patchout. arXiv preprint arXiv:2110.05069, 2021.
  31. 31.Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.
  32. 32.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  33. 33.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  34. 34.Ma, K. W., Wong, H. M., and Mak, C. M. A systematic review of human perceptual dimensions of sound: Meta-analysis of semantic differential method applications to indoor and outdoor sounds. Building and Environment, 133:123–150, 2018.
  35. 35.Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., and Sun, C. Attention bottlenecks for multimodal fusion. Advances in Neural Information Processing Systems, 34:14200–14213, 2021.
  36. 36.Niizumi, D., Takeuchi, D., Ohishi, Y., Harada, N., and Kashino, K. Byol for audio: Self-supervised learning for general-purpose audio representation. In 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1–8. IEEE, 2021.
  37. 37.Niizumi, D., Takeuchi, D., Ohishi, Y., Harada, N., and Kashino, K. Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation. arXiv preprint arXiv:2204.12260, 2022.
  38. 38.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 5206–5210. IEEE, 2015.
  39. 39.Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V. Specaugment: A simple data augmentation method for automatic speech recognition. arXiv preprint arXiv:1904.08779, 2019.
  40. 40.Patterson, K., Nestor, P. J., and Rogers, T. T. Where do you know what you know? the representation of semantic knowledge in the human brain. Nature reviews neuroscience, 8(12):976–987, 2007.
  41. 41.Peng, Z., Dong, L., Bao, H., Ye, Q., and Wei, F. BEiT v2: Masked image modeling with vector-quantized visual tokenizers. 2022.
  42. 42.Piczak, K. J. Esc: Dataset for environmental sound classification. In Proceedings of the 23rd ACM international conference on Multimedia, pp. 1015–1018, 2015.
  43. 43.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021.
  44. 44.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In International Conference on Machine Learning, pp. 8821–8831. PMLR, 2021.
  45. 45.Ravanelli, M. and Bengio, Y. Learning speaker representations with mutual information. arXiv preprint arXiv:1812.00271, 2018.
  46. 46.Reddy, C. K., Dubey, H., Koishida, K., Nair, A., Gopal, V., Cutler, R., Braun, S., Gamper, H., Aichner, R., and Srinivasan, S. Interspeech 2021 deep noise suppression challenge. arXiv preprint arXiv:2101.01902, 2021.
  47. 47.Saeed, A., Grangier, D., and Zeghidour, N. Contrastive learning of general-purpose audio representations. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3875–3879. IEEE, 2021.
  48. 48.Schneider, S., Baevski, A., Collobert, R., and Auli, M. wav2vec: Unsupervised pre-training for speech recognition. arXiv preprint arXiv:1904.05862, 2019.
  49. 49.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  50. 50.Tagliasacchi, M., Gfeller, B., de Chaumont Quitry, F., and Roblek, D. Pre-training audio representations with self-supervision. IEEE Signal Processing Letters, 27:600–604, 2020.
  51. 51.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  52. 52.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  53. 53.Verbitskiy, S., Berikov, V., and Vyshegorodtsev, V. Eranns: Efficient residual audio neural networks for audio pattern recognition. Pattern Recognition Letters, 161:38–44, 2022.
  54. 54.Wang, H., Ma, S., Dong, L., Huang, S., Zhang, D., and Wei, F. Deepnet: Scaling transformers to 1,000 layers. arXiv preprint arXiv:2203.00555, 2022a.
  55. 55.Wang, L. and Oord, A. v. d. Multi-format contrastive learning of audio representations. arXiv preprint arXiv:2103.06508, 2021.
  56. 56.Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022b.
  57. 57.Warden, P. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209, 2018.
  58. 58.wen Yang, S., Chi, P.-H., Chuang, Y.-S., Lai, C.-I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., Huang, T.-H., Tseng, W.-C., tik Lee, K., Liu, D.-R., Huang, Z., Dong, S., Li, S.-W., Watanabe, S., Mohamed, A., and yi Lee, H. SUPERB: Speech Processing Universal PERformance Benchmark. In Proc. Interspeech 2021, pp. 1194–1198, 2021. doi: 10.21437/Interspeech.2021-1775.
  59. 59.Wu, H.-H., Seetharaman, P., Kumar, K., and Bello, J. P. Wav2clip: Learning robust audio representations from clip. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4563–4567. IEEE, 2022.
  60. 60.Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., Feichtenhofer, C., et al. Masked autoencoders that listen. arXiv preprint arXiv:2207.06405, 2022.
  61. 61.Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627, 2021.
  62. 62.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  63. 63.Zhang, Y., Park, D. S., Han, W., Qin, J., Gulati, A., Shor, J., Jansen, A., Xu, Y., Huang, Y., Wang, S., et al. Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition. IEEE Journal of Selected Topics in Signal Processing, 16(6):1519–1532, 2022.

Citation

MLA
Chen, S., et al. “BEATs: Audio Pre-Training with Acoustic Tokenizers”. International Conference on Machine Learning, vol. 202, 2023, pp. 5178–93, https://proceedings.mlr.press/v202/chen23ag.html.
APA
Chen, S., Wu, Y., Wang, C., Liu, S., Tompkins, D., Chen, Z., Che, W., Yu, X., & Wei, F. (2023). BEATs: Audio Pre-Training with Acoustic Tokenizers. International Conference on Machine Learning, 202, 5178–5193. https://proceedings.mlr.press/v202/chen23ag.html
Chicago
Chen, S., Y. Wu, C. Wang, et al. 2023. “BEATs: Audio Pre-Training with Acoustic Tokenizers”. International Conference on Machine Learning 202: 5178–93. https://proceedings.mlr.press/v202/chen23ag.html.
Harvard
Chen, S. et al. (2023) “BEATs: Audio Pre-Training with Acoustic Tokenizers”, International Conference on Machine Learning. PMLR, pp. 5178–5193. Available at: https://proceedings.mlr.press/v202/chen23ag.html.
Vancouver
1. Chen S, Wu Y, Wang C, Liu S, Tompkins D, Chen Z, Che W, Yu X, Wei F (2023) BEATs: Audio Pre-Training with Acoustic Tokenizers. In: International Conference on Machine Learning. PMLR, pp 5178–5193

BibTeX

@InProceedings{pmlr-v202-chen23ag,
  title = 	 {{BEAT}s: Audio Pre-Training with Acoustic Tokenizers},
  author =       {Chen, Sanyuan and Wu, Yu and Wang, Chengyi and Liu, Shujie and Tompkins, Daniel and Chen, Zhuo and Che, Wanxiang and Yu, Xiangzhan and Wei, Furu},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {5178--5193},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/chen23ag/chen23ag.pdf},
  url = 	 {https://proceedings.mlr.press/v202/chen23ag.html},
  abstract = 	 {We introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/