BEATs: Audio Pre-Training with Acoustic Tokenizers
Sanyuan ChenYu WuChengyi WangShujie LiuDaniel TompkinsZhuo ChenWanxiang CheXiangzhan YuFuru Wei
Proposes an iterative self-supervised pre-training framework that jointly optimizes an acoustic tokenizer and an audio transformer via discrete label prediction, achieving state-of-the-art classification performance on AudioSet-2M and ESC-50 while using fewer parameters and less data than previous reconstruction-based methods.
Building powerful artificial intelligence models to understand general audio is challenging due to the diverse, unstructured nature of environmental sounds, background noise, and speech. While existing self-supervised audio models typically rely on reconstructing low-level acoustic details—such as raw spectrograms—this approach often retains background noise and fails to extract high-level semantic meaning. The article addresses this limitation by developing an iterative audio pre-training framework called BEATS (Bidirectional Encoder representation from Audio Transformers), which shifts the training objective from reconstructing low-level acoustic details to predicting high-level discrete acoustic tokens.
The framework pairs an acoustic tokenizer with a transformer-based audio encoder in an alternating, iterative training loop. Initially, a random-projection tokenizer creates baseline discrete labels. The audio encoder learns to predict these discrete tokens for masked audio segments. Once trained, the encoder acts as a teacher to guide and refine a new, self-distilled acoustic tokenizer through knowledge distillation. The newly refined tokenizer generates richer semantic labels to train the subsequent audio encoder, and the process repeats. The article pre-trained the framework on roughly two million clips from the AudioSet database and evaluated the resulting models across six standard audio and speech classification benchmarks.
The experimental findings show that BEATS consistently outperforms previous approaches while using fewer resources. First, single BEATS models established new state-of-the-art results on major benchmarks, achieving a 48.6% mean Average Precision (mAP) on the full AudioSet-2M dataset and a 25% relative error reduction on the ESC-50 environmental sound benchmark (reaching 98.1% accuracy). Second, BEATS outperformed previous leading models that required over three times as many parameters (90 million versus 304 million) or extensive pre-training on external visual datasets like ImageNet. Third, ensembling ten fine-tuned BEATS models achieved a record 50.6% mAP on AudioSet-2M without using any external data. Finally, visualization and disturbance analyses confirmed that the acoustic tokenizers successfully ignore background noise and reverberation, isolating meaningful semantic content far better than reconstruction-based methods.
These findings indicate that discrete label prediction enables more semantically focused and computationally efficient audio representation learning. By matching or exceeding larger systems with significantly smaller model architectures, this approach lowers deployment costs and reduces dependence on massive supervised labels or cross-modal pre-training data. Furthermore, using discrete tokens aligns audio pre-training methods with those used in natural language, computer vision, and speech processing, offering a viable pathway toward building unified multimodal foundation models.
Stakeholders and engineering teams should consider adopting the released BEATS pre-trained models and acoustic tokenizers to optimize sound classification workflows, lower labeling overhead, and accelerate downstream application development. For future work, development efforts should focus on scaling the architecture beyond 90 million parameters, expanding pre-training data beyond AudioSet to broader audio domains, and integrating the tokenizer framework into multimodal systems combining audio, text, and vision. While the reported performance gains are statistically robust across multiple standard benchmarks, the main practical trade-off is the linear increase in training compute required during the iterative tokenizer-model optimization cycles.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT established the masked prediction pre-training paradigm using discrete visual tokenizers, providing the conceptual foundation for BEATs' acoustic tokenizer and discrete label prediction framework.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). HuBERT introduced iterative self-supervised representation learning via masked prediction of discrete hidden cluster units, which directly motivates the iterative tokenizer and SSL optimization design in BEATs.
- Paper: AST: Audio Spectrogram Transformer, Yuan Gong et al. (2021). AST pioneered the use of pure Vision Transformer architectures on spectrogram patches for audio classification, establishing the baseline neural backbone used in BEATs.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). wav2vec 2.0 provides essential background on self-supervised audio pre-training via latent quantization and masked temporal modeling.
- Paper: WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, Sanyuan Chen et al. (2021). WavLM demonstrates large-scale masked self-supervised speech modeling, setting the benchmark for full-stack acoustic pre-training that BEATs generalizes to broader audio classification.
- Paper: SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition, Daniel S. Park et al. (2019). SpecAugment introduced time and frequency spectrogram masking techniques that form the core input corruption strategies in BEATs' pre-training and fine-tuning.
- Paper: CNN architectures for large-scale audio classification, Shawn Hershey et al. (2016). This paper establishes the large-scale YouTube/AudioSet audio classification setup and evaluation protocol upon which BEATs benchmarks its state-of-the-art representations.
- Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). CAV-MAE extends masked autoencoding and self-supervised acoustic representations into multimodal audio-visual classification and cross-modal retrieval on AudioSet.
- Paper: MAViL: Masked Audio-Video Learners, Po-Yao Huang et al. (2023). MAViL builds on self-supervised audio-visual learning by integrating masked representations with iterative student-teacher training across AudioSet benchmarks.
- Paper: Pengi: An Audio Language Model for Audio Tasks, Soham Deshmukh et al. (2023). Pengi builds upon general-purpose audio representation models like BEATs to unify audio understanding tasks into a generative audio-language modeling framework.
- Paper: AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension, Qian Yang et al. (2024). AIR-Bench provides an extensive downstream benchmark for evaluating advanced audio and audio-language models across diverse perceptual and conversational sound tasks.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). MMAU extends the evaluation of foundational audio understanding representations to complex multi-task reasoning and expert domain comprehension.
