VoxCeleb2: Deep Speaker Recognition

Joon Son ChungArsha NagraniAndrew Zisserman

article2018Interspeech2,860 citations

Introduces the massive VoxCeleb2 dataset containing over one million utterances across six thousand speakers alongside deep convolutional neural network architectures that dramatically improve speaker recognition in unconstrained, noisy environments.

Listen

Speaker recognition remains difficult in noisy, real-world settings because compact voice representations that work reliably across accents, ages, and recording conditions are hard to produce. Large public datasets have been scarce, limiting progress compared to fields such as face recognition.

The article set out to create a much larger training resource and to test whether deep convolutional networks trained on it could improve verification performance under unconstrained conditions.

Researchers built VoxCeleb2 with an automated pipeline that downloads YouTube videos, tracks faces, verifies identity and active speech, and removes duplicates. The resulting set contains more than one million utterances from over six thousand speakers and is more than five times larger than the prior VoxCeleb1 collection. Models based on modified ResNet architectures were first trained for speaker identification, then fine-tuned with contrastive loss on spectrograms to produce 512-dimensional embeddings.

The best model, a ResNet-50 trained on VoxCeleb2, reached an equal error rate of 3.95 percent on the original VoxCeleb1 test set, compared with 7.8 percent for the previous best system. On two new, larger test sets drawn from the full VoxCeleb1 collection the same model achieved 4.42 percent and 7.33 percent equal error rates. Performance improved steadily with network depth and with the larger training set.

These results show that scale and modern residual architectures together produce more robust speaker embeddings that can be stored compactly and reused for verification, clustering, or diarisation. The gains matter for any application that must identify speakers from everyday audio without controlled recording conditions.

The authors recommend adopting the two new VoxCeleb1 evaluation protocols as standard benchmarks alongside existing sets such as SITW. They also release the full VoxCeleb2 dataset to support further research.

The main limitations are reliance on an automated collection pipeline that may still contain a small number of label errors and the fact that all reported results come from a single source of celebrity interview footage. Confidence in the performance ordering is high because consistent improvements appear across multiple architectures and test conditions, yet absolute numbers could shift on other domains or languages.

Cover for VoxCeleb2: Deep Speaker Recognition

Abstract

The objective of this paper is speaker recognition under noisy and unconstrained conditions.

We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media. Using a fully automated pipeline, we curate VoxCeleb2 which contains over a million utterances from over 6,000 speakers. This is several times larger than any publicly available speaker recognition dataset.

Second, we develop and compare Convolutional Neural Network (CNN) models and training strategies that can effectively recognise identities from voice under various conditions. The models trained on the VoxCeleb2 dataset surpass the performance of previous works on a benchmark dataset by a significant margin.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 The VoxCeleb2 Dataset
  • 3.1 Description
  • 3.2 Collection Pipeline
  • 4 VGGVox
  • 4.1 Evaluation
  • 4.2 Trunk architectures
  • 4.3 Training Loss strategies
  • 4.4 Test time augmentation
  • 4.5 Implementation Details
  • 5 Results
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — The VoxCeleb2 Audio-Visual Speaker Dataset

    data/table

    The VoxCeleb2 dataset is a large-scale audio-visual dataset curated automatically from YouTube videos for speaker recognition in unconstrained, noisy conditions. It contains over 1.12 million utterances from 6,112 persons of interest (POIs), spanning 145 nationalities with 61% male speakers. The development set contains no identity overlap with the VoxCeleb1 or SITW datasets.

    Dataset Metric VoxCeleb1 VoxCeleb2
    # of POIs 1,251 6,112
    # of male POIs 690 3,761
    # of videos 22,496 150,480
    # of hours 352 2,442
    # of utterances 153,516 1,128,246
    Avg # of videos per POI 18 25
    Avg # of utterances per POI 116 185
    Avg length of utterances (s) 8.2 7.8
    Subset Development Test Total
    # of POIs 5,994 118 6,112
    # of videos 145,569 4,911 150,480
    # of utterances 1,092,009 36,237 1,128,246
  2. Knowl 2 — Automated Multi-Stage Audio-Visual Curation Pipeline

    model/method

    The VoxCeleb2 dataset is curated using a fully automated computer vision and audio pipeline consisting of seven stages:

    1. Candidate Person of Interest (POI) Selection: Candidate identities are sourced from the VGGFace2 dataset (over 9,000 identities across diverse professions and ethnicities). Any POIs overlapping with VoxCeleb1 or SITW are removed from the development split.
    2. Video Downloading: Up to 100 videos per POI are queried and downloaded automatically from YouTube by searching for <POI name> interview.
    3. Face Detection and Tracking: A Single Shot MultiBox Detector (SSD) detects faces across multi-angle and profile views on each video frame. Continuous face tracks are constructed based on region-of-interest (ROI) bounding box overlap.
    4. Face Verification: A ResNet-50 network pre-trained on VGGFace2 classifies face tracks to verify whether they depict the targeted POI.
    5. Active Speaker Verification: A multi-view adaptation of SyncNet (a two-stream CNN) determines whether the visible face is actively speaking by estimating the temporal correlation between mouth motion and the audio track, filtering out dubbing and voice-overs.
    6. Audio Duplicate Removal: Each speech segment is mapped to a 1024-dimensional feature vector using a deep audio feature extractor. If the Euclidean distance between any two segment vectors from the same POI is below a conservative threshold of 0.10.1, the segment is tagged as a duplicate (or near-duplicate trimmed version) and discarded.
    7. Demographic Labeling: Nationality (country of citizenship) labels are scraped from Wikipedia, successfully labeling all but 428 speakers across 145 countries.
  3. Knowl 3 — Spectrogram-Adapted ResNet Trunk Architectures

    model/method

    VGGVox adapts residual networks (ResNet-34 and ResNet-50) to process short-term magnitude spectrograms (dimensions 512×300512 \times 300 for 3 seconds of audio computed with a 25ms Hamming window and 10ms step). Batch normalization is applied before rectified linear unit (ReLU) activations.

    To ensure temporal invariance while preserving frequency information, the standard 2D global average pooling is replaced by an fc1 layer with support 9×19 \times 1 in the frequency domain, followed by a 1D temporal average pooling layer (pool_time) across time dimension NN. The network concludes with a linear projection layer fc2 (1×11 \times 1) mapping to the 5,994 POI classes during pre-training or a 512-dimensional embedding space during metric learning.

    Layer Name ResNet-34 ResNet-50
    conv1 7×7,647 \times 7, 64, stride 2 7×7,647 \times 7, 64, stride 2
    pool1 3×33 \times 3, max pool, stride 2 3×33 \times 3, max pool, stride 2
    conv2_x [3×3,643×3,64]×3\begin{bmatrix} 3 \times 3, 64 \\ 3 \times 3, 64 \end{bmatrix} \times 3 [1×1,643×3,641×1,256]×3\begin{bmatrix} 1 \times 1, 64 \\ 3 \times 3, 64 \\ 1 \times 1, 256 \end{bmatrix} \times 3
    conv3_x [3×3,1283×3,128]×4\begin{bmatrix} 3 \times 3, 128 \\ 3 \times 3, 128 \end{bmatrix} \times 4 [1×1,1283×3,1281×1,512]×4\begin{bmatrix} 1 \times 1, 128 \\ 3 \times 3, 128 \\ 1 \times 1, 512 \end{bmatrix} \times 4
    conv4_x [3×3,2563×3,256]×6\begin{bmatrix} 3 \times 3, 256 \\ 3 \times 3, 256 \end{bmatrix} \times 6 [1×1,2563×3,2561×1,1024]×6\begin{bmatrix} 1 \times 1, 256 \\ 3 \times 3, 256 \\ 1 \times 1, 1024 \end{bmatrix} \times 6
    conv5_x [3×3,5123×3,512]×3\begin{bmatrix} 3 \times 3, 512 \\ 3 \times 3, 512 \end{bmatrix} \times 3 [1×1,5123×3,5121×1,2048]×3\begin{bmatrix} 1 \times 1, 512 \\ 3 \times 3, 512 \\ 1 \times 1, 2048 \end{bmatrix} \times 3
    fc1 9×1,5129 \times 1, 512, stride 1 9×1,20489 \times 1, 2048, stride 1
    pool_time 1×N1 \times N, avg pool, stride 1 1×N1 \times N, avg pool, stride 1
    fc2 1×1,59941 \times 1, 5994 1×1,59941 \times 1, 5994
  4. Knowl 4 — Two-Stage Embedding Training with Offline Hard Negative Mining

    model/method

    Training VGGVox speaker embeddings proceeds in two sequential stages to avoid poor local minima:

    1. Softmax Pre-Training: The CNN trunk architecture is trained on identification using cross-entropy loss over the 5,994 POIs in the VoxCeleb2 development set, providing stable initial representations.
    2. Contrastive Loss Fine-Tuning: The final 5,994-way classification layer is replaced with a 512-dimensional fully connected projection layer. The network is fine-tuned using a contrastive loss with margin α\alpha, which minimizes the Euclidean distance between positive (same-speaker) embedding pairs while penalizing negative (different-speaker) pairs with distance below α\alpha.

    To handle quadratic pair growth, an offline hard negative mining strategy is used: negative pairs are generated offline and filtered to retain only the top 1% hardest negatives (those with smallest Euclidean distance). Hard positive mining is deliberately avoided because potential label noise in automated face verification would create false positive pairs and destabilize optimization.

  5. Knowl 5 — Speaker Verification Detection Cost Function

    equation

    The performance of speaker verification systems is evaluated using Equal Error Rate (EER) and the normalized minimum detection cost function CdetC_{\text{det}}:

    Cdet=Cmiss×Pmiss×Ptar+Cfa×Pfa×(1−Ptar)C_{\text{det}} = C_{\text{miss}} \times P_{\text{miss}} \times P_{\text{tar}} + C_{\text{fa}} \times P_{\text{fa}} \times (1 - P_{\text{tar}})

    where:

    • Ptar∈[0,1]P_{\text{tar}} \in [0, 1] is the prior probability that a test segment belongs to the target speaker, set to Ptar=0.01P_{\text{tar}} = 0.01.
    • CmissC_{\text{miss}} is the cost penalty assigned to a false rejection (miss), set to 1.01.0.
    • CfaC_{\text{fa}} is the cost penalty assigned to a false acceptance (false alarm), set to 1.01.0.
    • Pmiss∈[0,1]P_{\text{miss}} \in [0, 1] is the empirical probability of a miss at a chosen decision threshold.
    • Pfa∈[0,1]P_{\text{fa}} \in [0, 1] is the empirical probability of a false alarm at the same threshold.
  6. Knowl 6 — Test-Time Augmentation Strategies for Speaker Verification

    model/method

    Given two variable-length test utterances to compare, speaker similarity is evaluated using one of three test-time protocols:

    • Method (1) Full-Utterance Average Pooling: The full spectrogram of each utterance is passed through the network, dynamically adjusting the temporal pooling layer (pool_time) size to match the input length, generating one embedding per utterance.
    • Method (2) Mean Crop Feature Pooling: Ten 3-second temporal crops are sampled from each utterance. The 512-dimensional embeddings of the 10 crops are averaged into a single mean embedding per utterance, and the Euclidean distance between the two mean embeddings is computed.
    • Method (3) Pairwise Crop Distance Averaging: Ten 3-second temporal crops are sampled from each utterance. Pairwise Euclidean distances are computed for all 10×10=10010 \times 10 = 100 crop combinations across the two utterances, and the mean of these 100 distances is used as the final verification score.
  7. Knowl 7 — VoxCeleb1-E and VoxCeleb1-H Verification Benchmarks

    definition

    To evaluate speaker verification across a broader demographic distribution than the original 40-speaker VoxCeleb1 test set, two evaluation protocols are defined over the entire 1,251 speakers of VoxCeleb1:

    • VoxCeleb1-E (Entire): An evaluation protocol comprising 581,480 randomly sampled trial pairs drawn from all 1,251 speakers of the VoxCeleb1 dataset.
    • VoxCeleb1-H (Hard): An evaluation protocol consisting of 552,536 trial pairs where both segments in each pair belong to individuals of the same nationality and same gender. The pairs are drawn from 18 distinct nationality-gender groups that contain at least 5 individual speakers (with the 'USA-Male' demographic being the largest).
  8. Knowl 8 — Speaker Verification Performance on the Original VoxCeleb1 Test Set

    empirical result

    Scaling training data from VoxCeleb1 to VoxCeleb2 and increasing model depth from VGG-M to ResNet-50 substantially improves verification performance on the original VoxCeleb1 test set. Test-time crop distance averaging (Method 3) provides consistent additional gains.

    Models Trained on Cdetmin⁡C_{\text{det}}^{\min} EER (%)
    I-vectors + PLDA (1) VoxCeleb1 0.73 8.8
    VGG-M (Softmax) VoxCeleb1 0.75 10.2
    VGG-M (1) VoxCeleb1 0.71 7.8
    VGG-M (1) VoxCeleb2 0.609 5.94
    ResNet-34 (1) VoxCeleb2 0.543 5.04
    ResNet-34 (2) VoxCeleb2 0.553 5.11
    ResNet-34 (3) VoxCeleb2 0.549 4.83
    ResNet-50 (1) VoxCeleb2 0.449 4.19
    ResNet-50 (2) VoxCeleb2 0.454 4.43
    ResNet-50 (3) VoxCeleb2 0.429 3.95

    The numbers in parentheses refer to the test-time augmentation method used (1: full-utterance average pooling; 2: mean crop feature pooling; 3: pairwise crop distance averaging).

  9. Knowl 9 — Speaker Verification Performance on VoxCeleb1-E and VoxCeleb1-H Test Sets

    empirical result

    Evaluating the ResNet-50 architecture (using test augmentation Method 3) on the extended VoxCeleb1 test protocols yields the following verification performance:

    Model Tested on Cdetmin⁡C_{\text{det}}^{\min} EER (%)
    ResNet-50 (3) VoxCeleb1-E 0.524 4.42
    ResNet-50 (3) VoxCeleb1-H 0.673 7.33

    The higher EER on VoxCeleb1-H (7.33% vs 4.42%) highlights that verifying speakers within the same gender and nationality presents a significantly more difficult biometric challenge due to the absence of confounding acoustic cues from accent and gender.

Coverage note — All primary contributions—including dataset statistics, automated curation pipeline, ResNet spectrogram adaptations, two-stage metric learning with offline hard negative mining, test-time augmentations, evaluation cost formulations, and empirical verification benchmarks—have been converted into knowls. Standard baseline descriptions and routine software/GPU training environments were omitted as they are not core novel contributions.

References

  1. 1.F. Schroff, D. Kalenichenko, and J. Philbin, ‘Facenet: A unified embedding for face recognition and clustering,’ in Proc. CVPR, 2015.
  2. 2.Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, ‘Deepface: Closing the gap to human-level performance in face verification,’ in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 1701–1708.
  3. 3.O. M. Parkhi, A. Vedaldi, and A. Zisserman, ‘Deep face recognition,’ in Proc. BMVC., 2015.
  4. 4.Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman, ‘VGGFace2: a dataset for recognising faces across pose and age,’ arXiv preprint arXiv:1710.08092, 2017.
  5. 5.I. Kemelmacher-Shlizerman, S. M. Seitz, D. Miller, and E. Brossard, ‘The megaface benchmark: 1 million faces for recognition at scale,’ in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4873–4882.
  6. 6.Y. Guo, L. Zhang, Y. Hu, X. He, and J. Gao, ‘MS-Celeb-1M: A dataset and benchmark for large-scale face recognition,’ in European Conference on Computer Vision. Springer, 2016, pp. 87–102.
  7. 7.A. Nagrani, J. S. Chung, and A. Zisserman, ‘VoxCeleb: a large-scale speaker identification dataset,’ in INTERSPEECH, 2017.
  8. 8.M. McLaren, L. Ferrer, D. Castan, and A. Lawson, ‘The speakers in the wild (SITW) speaker recognition database,’ in INTERSPEECH, 2016.
  9. 9.J. S. Chung, A. Jamaludin, and A. Zisserman, ‘You said that?’ in Proc. BMVC., 2017.
  10. 10.T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, ‘Audio-driven facial animation by joint end-to-end learning of pose and emotion,’ ACM Transactions on Graphics (TOG), vol. 36, no. 4, p. 94, 2017.
  11. 11.T. Afouras, J. S. Chung, and A. Zisserman, ‘The conversation: Deep audio-visual speech enhancement,’ in arXiv:1804.04121, 2018.
  12. 12.A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, ‘Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,’ arXiv preprint arXiv:1804.03619, 2018.
  13. 13.A. Nagrani, S. Albanie, and A. Zisserman, ‘Seeing voices and hearing faces: Cross-modal biometric matching,’ in IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  14. 14.A. Nagrani, S. Albanie, and A. Zisserman, ‘Learnable pins: Cross-modal embeddings for person identity,’ arXiv preprint arXiv:1805.00833, 2018.
  15. 15.K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, ‘Return of the devil in the details: Delving deep into convolutional nets,’ in Proc. BMVC., 2014.
  16. 16.K. He, X. Zhang, S. Ren, and J. Sun, ‘Deep residual learning for image recognition,’ arXiv preprint arXiv:1512.03385, 2015.
  17. 17.N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, ‘Front-end factor analysis for speaker verification,’ IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 4, pp. 788–798, 2011.
  18. 18.P. Matejka, O. Glembek, F. Castaldo, M. J. Alam, O. Plchot, P. Kenny, L. Burget, and J. Cernocky, ‘Full-covariance ubm and heavy-tailed plda in i-vector speaker verification,’ in Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on. IEEE, 2011, pp. 4828–4831.
  19. 19.S. Cumani, O. Plchot, and P. Laface, ‘Probabilistic linear discriminant analysis of i-vector posterior distributions,’ in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013, pp. 7644–7648.
  20. 20.J. H. Hansen and T. Hasan, ‘Speaker recognition by machines and humans: A tutorial review,’ IEEE Signal processing magazine, vol. 32, no. 6, pp. 74–99, 2015.
  21. 21.E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, ‘Deep neural networks for small footprint text-dependent speaker verification,’ in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 4052–4056.
  22. 22.Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, ‘A novel scheme for speaker recognition using a phonetically-aware deep neural network,’ in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 1695–1699.
  23. 23.S. H. Ghalehjegh and R. C. Rose, ‘Deep bottleneck features for i-vector based text-independent speaker verification,’ in Automatic Speech Recognition and Understanding (ASRU), 2015 IEEE Workshop on. IEEE, 2015, pp. 555–560.
  24. 24.D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur, ‘Deep neural network embeddings for text-independent speaker verification,’ Proc. Interspeech 2017, pp. 999–1003, 2017.
  25. 25.D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, ‘X-vectors: Robust dnn embeddings for speaker recognition,’ ICASSP, Calgary, 2018.
  26. 26.D. Chen, S. Tsai, V. Chandrasekhar, G. Takacs, H. Chen, R. Vedantham, R. Grzeszczuk, and B. Girod, ‘Residual enhanced visual vectors for on-device image matching,’ in Asilomar, 2011.
  27. 27.S. H. Yella, A. Stolcke, and M. Slaney, ‘Artificial neural network features for speaker diarization,’ in Spoken Language Technology Workshop (SLT), 2014 IEEE. IEEE, 2014, pp. 402–406.
  28. 28.C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu, ‘Deep speaker: an end-to-end neural speaker embedding system,’ arXiv preprint arXiv:1705.02304, 2017.
  29. 29.D. van der Vloed, J. Bouten, and D. A. van Leeuwen, ‘NFI-FRITS: a forensic speaker recognition database and some first experiments,’ in The Speaker and Language Recognition Workshop, 2014.
  30. 30.J. Hennebert, H. Melin, D. Petrovska, and D. Genoud, ‘POLYCOST: a telephone-speech database for speaker recognition,’ Speech communication, vol. 31, no. 2, pp. 265–270, 2000.
  31. 31.J. B. Millar, J. P. Vonwiller, J. M. Harrington, and P. J. Dermody, ‘The Australian national database of spoken language,’ in Proc. ICASSP, vol. 1. IEEE, 1994, pp. I–97.
  32. 32.J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, ‘DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,’ NASA STI/Recon technical report, vol. 93, 1993.
  33. 33.W. M. Fisher, G. R. Doddington, and K. M. Goudie-Marshall, ‘The DARPA speech recognition research database: specifications and status,’ in Proc. DARPA Workshop on speech recognition, 1986, pp. 93–99.
  34. 34.C. S. Greenberg, ‘The NIST year 2012 speaker recognition evaluation plan,’ NIST, Technical Report, 2012.
  35. 35.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, ‘Ssd: Single shot multibox detector,’ in Proc. ECCV. Springer, 2016, pp. 21–37.
  36. 36.J. S. Chung and A. Zisserman, ‘Lip reading in profile,’ in Proc. BMVC., 2017.
  37. 37.J. S. Chung and A. Zisserman, ‘Out of time: automated lip sync in the wild,’ in Workshop on Multi-view Lip-reading, ACCV, 2016.
  38. 38.J. S. Chung and A. Zisserman, ‘Learning to lip read words by watching videos,’ CVIU, 2018.
  39. 39.S. Chopra, R. Hadsell, and Y. LeCun, ‘Learning a similarity metric discriminatively, with application to face verification,’ in Proc. CVPR, vol. 1. IEEE, 2005, pp. 539–546.
  40. 40.R. Hadsell, S. Chopra, and Y. LeCun, ‘Dimensionality reduction by learning an invariant mapping,’ in CVPR, vol. 2. IEEE, 2006, pp. 1735–1742.
  41. 41.A. Hermans, L. Beyer, and B. Leibe, ‘In defense of the triplet loss for person re-identification,’ arXiv preprint arXiv:1703.07737, 2017.
  42. 42.K.-K. Sung, ‘Learning and example selection for object and pattern detection,’ Ph.D. dissertation, 1996.
  43. 43.H. O. Song, Y. Xiang, S. Jegelka, and S. Savarese, ‘Deep metric learning via lifted structured feature embedding,’ in Computer Vision and Pattern Recognition (CVPR), 2016 IEEE Conference on. IEEE, 2016, pp. 4004–4012.
  44. 44.A. Vedaldi and K. Lenc, ‘Matconvnet – convolutional neural networks for matlab,’ CoRR, vol. abs/1412.4564, 2014.

Citation

MLA
Chung, J. S., et al. “VoxCeleb2: Deep Speaker Recognition”. Interspeech 2018, 2018, pp. 1086–90, https://doi.org/10.21437/Interspeech.2018-1929.
APA
Chung, J. S., Nagrani, A., & Zisserman, A. (2018). VoxCeleb2: Deep Speaker Recognition. Interspeech 2018, 1086–1090. https://doi.org/10.21437/Interspeech.2018-1929
Chicago
Chung, J. S., A. Nagrani, and A. Zisserman. 2018. “VoxCeleb2: Deep Speaker Recognition”. Interspeech 2018, 1086–90. https://doi.org/10.21437/Interspeech.2018-1929.
Harvard
Chung, J.S., Nagrani, A. and Zisserman, A. (2018) “VoxCeleb2: Deep Speaker Recognition”, Interspeech 2018. ISCA, pp. 1086–1090. Available at: https://doi.org/10.21437/Interspeech.2018-1929.
Vancouver
1. Chung JS, Nagrani A, Zisserman A (2018) VoxCeleb2: Deep Speaker Recognition. In: Interspeech 2018. ISCA, pp 1086–1090

BibTeX

@inproceedings{Chung_2018, series={interspeech_2018}, title={VoxCeleb2: Deep Speaker Recognition}, url={http://dx.doi.org/10.21437/Interspeech.2018-1929}, DOI={10.21437/interspeech.2018-1929}, booktitle={Interspeech 2018}, publisher={ISCA}, author={Chung, Joon Son and Nagrani, Arsha and Zisserman, Andrew}, year={2018}, month=Sept, pages={1086–1090}, collection={interspeech_2018} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF