AST: Audio Spectrogram Transformer
Yuan GongYu-An ChungJames R. Glass
Introduces the Audio Spectrogram Transformer, the first convolution-free, purely attention-based architecture for audio classification, proving that self-attention alone outperforms standard convolutional networks across major audio and speech benchmarks.
Audio classification systems power critical applications such as voice command recognition, sound event detection, and audio monitoring. For over a decade, standard architectures have relied heavily on convolutional neural networks, often paired with attention mechanisms, based on the assumption that convolutional layers are essential for processing visual representations of audio known as spectrograms. However, maintaining customized convolutional designs across diverse audio tasks creates architectural complexity and requires significant domain-specific engineering.
The article aims to determine whether convolutional components are truly indispensable for audio analysis by introducing and evaluating the Audio Spectrogram Transformer, the first fully attention-based, convolution-free architecture designed for audio classification. The authors demonstrate that an unmodified transformer encoder applied directly to spectrogram patches, combined with cross-modal knowledge transfer from vision models, can effectively classify variable-length audio streams across multiple domains.
The authors conducted rigorous experimental evaluations across three primary benchmark datasets: AudioSet (weakly labeled audio events spanning 2 million YouTube clips), ESC-50 (environmental sound classification across 50 classes), and Speech Commands V2 (35-class keyword spotting across 105,829 recordings). To overcome the data-hungry nature of pure transformer models, the team repurposed off-the-shelf vision transformers pretrained on large-scale image databases. They adapted positional embeddings and channel structures to accommodate audio spectrograms of variable length (from 1 to 10 seconds) without altering the core underlying neural network architecture.
The findings show that the proposed attention-based model systematically outperforms existing state-of-the-art hybrid and convolutional models. On AudioSet, the model achieved a record 0.485 mean average precision in an ensemble configuration and 0.459 in a single-model configuration, exceeding previous hybrid baselines while requiring only 5 training epochs compared to 30 for earlier networks. On ESC-50, the model attained 95.6% accuracy with audio pretraining and 88.7% without, surpassing prior benchmarks. On Speech Commands V2, it set a new top accuracy of 98.11% without requiring supplementary audio pretraining. Ablation studies revealed that transferring pretrained vision representations is the most crucial performance driver, more than doubling baseline capability on smaller training subsets.
These results demonstrate that convolutional layers are not required for high-accuracy audio modeling. The findings carry significant operational implications: organizations can replace complex, specialized convolutional pipelines with standard, unified transformer architectures. This reduces custom development timelines, lowers hyperparameter tuning overhead across different audio lengths, and accelerates training cycles by up to 80% while establishing new performance highs.
Engineering and research teams should adopt standard transformer architectures and leverage existing vision pretraining pipelines for speech and sound recognition applications instead of designing custom convolutional networks. Teams can adjust patch overlap based on computational budgets, as higher overlap yields marginally better accuracy at the expense of quadratic computational scaling. Further research should explore optimal pretraining strategies tailored directly to rectangular spectrogram representations.
Readers should note that the model relies heavily on cross-modal transfer from image pretraining; training purely attention-based audio transformers from scratch on small datasets remains suboptimal. Additionally, higher patch overlap increases memory requirements during training. Despite these boundary conditions, the empirical evidence across multiple distinct benchmarks provides high confidence in the framework's broad utility.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces the foundational Transformer architecture and multi-head self-attention mechanism that AST adapts for 2D audio spectrogram classification.
- Paper: CNN architectures for large-scale audio classification, Shawn Hershey et al. (2016). Establishes standard deep convolutional network architectures and benchmarking practices on large-scale spectrogram audio classification that AST seeks to replace with attention.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Provides the Vision Transformer training protocols and ImageNet pre-training transfer strategies directly utilized by AST to initialize spectrogram patch transformers.
- Paper: Conformer: Convolution-augmented Transformer for Speech Recognition, Anmol Gulati et al. (2020). Presents the hybrid convolution-transformer paradigm for speech and audio sequences, establishing the baseline tension between local convolution and global attention that AST evaluates.
- Paper: Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification, Justin Salamon et al. (2016). Demonstrates core spectrogram representation paradigms and data augmentation methods for environmental sound classification evaluated in AST.
- Paper: A Survey on Vision Transformer, Kai Han et al. (2020). Surveys the patch-based Vision Transformer methodology that AST translates from 2D natural images to 2D time-frequency audio representations.
- Paper: Perceiver: General Perception with Iterative Attention, Andrew Jaegle et al. (2021). Generalizes pure-attention architectures beyond discrete spectrogram patching to iterative latent attention across arbitrary raw multimodal signals including audio and video.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Systematically reviews how attention mechanisms and patch tokenization methods, such as those pioneered in AST, extend to continuous time-series modeling and classification.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). Examines whether modernized pure convolutional networks can challenge and reclaim the performance advantages established by pure transformer models like AST and ViT.
