VoxCeleb2: Deep Speaker Recognition
Joon Son ChungArsha NagraniAndrew Zisserman
Introduces the massive VoxCeleb2 dataset containing over one million utterances across six thousand speakers alongside deep convolutional neural network architectures that dramatically improve speaker recognition in unconstrained, noisy environments.
Speaker recognition remains difficult in noisy, real-world settings because compact voice representations that work reliably across accents, ages, and recording conditions are hard to produce. Large public datasets have been scarce, limiting progress compared to fields such as face recognition.
The article set out to create a much larger training resource and to test whether deep convolutional networks trained on it could improve verification performance under unconstrained conditions.
Researchers built VoxCeleb2 with an automated pipeline that downloads YouTube videos, tracks faces, verifies identity and active speech, and removes duplicates. The resulting set contains more than one million utterances from over six thousand speakers and is more than five times larger than the prior VoxCeleb1 collection. Models based on modified ResNet architectures were first trained for speaker identification, then fine-tuned with contrastive loss on spectrograms to produce 512-dimensional embeddings.
The best model, a ResNet-50 trained on VoxCeleb2, reached an equal error rate of 3.95 percent on the original VoxCeleb1 test set, compared with 7.8 percent for the previous best system. On two new, larger test sets drawn from the full VoxCeleb1 collection the same model achieved 4.42 percent and 7.33 percent equal error rates. Performance improved steadily with network depth and with the larger training set.
These results show that scale and modern residual architectures together produce more robust speaker embeddings that can be stored compactly and reused for verification, clustering, or diarisation. The gains matter for any application that must identify speakers from everyday audio without controlled recording conditions.
The authors recommend adopting the two new VoxCeleb1 evaluation protocols as standard benchmarks alongside existing sets such as SITW. They also release the full VoxCeleb2 dataset to support further research.
The main limitations are reliance on an automated collection pipeline that may still contain a small number of label errors and the fact that all reported results come from a single source of celebrity interview footage. Confidence in the performance ordering is high because consistent improvements appear across multiple architectures and test conditions, yet absolute numbers could shift on other domains or languages.
- Paper: Deep Face Recognition, Omkar M. Parkhi et al. (2015). Introduces the scalable web-scraping curation pipeline and metric learning embedding strategies directly adapted by the VoxCeleb series to construct large-scale identity recognition datasets.
- Paper: VGGFace2: A Dataset for Recognising Faces across Pose and Age, Qiong Cao et al. (2017). Presents the large-scale VGGFace2 dataset and deep ResNet identity architectures that served as the foundational visual counterpart and methodological inspiration for VoxCeleb2.
- Paper: Learning a similarity metric discriminatively, with application to face verification, Sumit Chopra et al. (2005). Establishes Siamese neural networks and contrastive loss optimization for identity verification, forming the core metric-learning framework used to fine-tune VoxCeleb2 speaker embeddings.
- Paper: CNN architectures for large-scale audio classification, Shawn Hershey et al. (2016). Demonstrates how 2D convolutional networks such as ResNet can effectively extract robust representations directly from audio spectrograms at scale.
- Paper: A Discriminative Feature Learning Approach for Deep Face Recognition, Yandong Wen et al. (2016). Introduces discriminative metric learning principles for learning compact intra-class feature embeddings for open-set identification and verification.
- Paper: WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, Sanyuan Chen et al. (2021). Leverages large-scale pre-training across speech benchmarks and directly utilizes VoxCeleb evaluation protocols to establish state-of-the-art speaker verification performance.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). Advances speech representation learning through masked acoustic unit prediction, providing an alternative self-supervised paradigm to VoxCeleb2's supervised speaker embedding approach.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Develops a self-supervised framework on raw audio that extends the representational power of speech models beyond supervised speaker embeddings.
- Paper: Conformer: Convolution-augmented Transformer for Speech Recognition, Anmol Gulati et al. (2020). Combines convolution with self-attention to capture both local spectrogram features and long-range temporal contexts in speech processing.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). Generalizes cross-modal representation learning by modeling unaligned audio and visual sequence interactions.
