Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation

Martijn BarteldsNay SanBradley McDonnellDan JurafskyMartijn Wieling

article2023ACL78 citations

Demonstrates how self-training and text-to-speech data augmentation can reduce word error rates by up to 25.5% when fine-tuning speech recognition models on as little as 24 minutes of transcribed audio across four diverse minority languages.

Listen

Recent advances in automated speech recognition have largely benefited major world languages, leaving regional, minority, and dialectal communities digitally underserved due to a lack of large transcribed audio datasets. Generating manual transcriptions for these languages requires substantial human effort and domain expertise, making data scarcity a central barrier to deploying language technologies in real-world documentation and preservation efforts.

The article demonstrates and evaluates whether practical data augmentation techniques can cost-effectively improve speech recognition performance for severely resource-scarce languages. Specifically, the authors assess the impact of self-training across four typologically diverse minority languages and examine synthetic speech generation using text-to-speech technology where existing systems allow.

To conduct this evaluation, the researchers analyzed four minority languages: Gronings, West-Frisian, Besemah, and Nasal. For each language, the total pool of human-transcribed speech was restricted to four hours. The team fine-tuned a standard multilingual pre-trained foundation model across varying training sizes (from 24 to 192 minutes) on a single high-performance graphics processing unit. They tested self-training by using an initial model to transcribe unannotated audio and retraining subsequent models on the combined data. Additionally, for Gronings, they augmented models using synthetic speech generated from text via an existing text-to-speech system.

The findings show that self-training consistently improves accuracy across all four languages, reducing word error rates by 6.3% to 13.9% relative to baseline models trained on just 24 minutes of transcribed audio. Iterative self-training with expanded unannotated audio delivered relative error reductions of up to 20.5% for Gronings, matching the performance of a model trained on double the volume of human-transcribed data. Furthermore, augmenting training with text-to-speech data yielded even larger improvements, achieving up to a 25.5% error reduction on out-of-domain evaluation and performing on par with models using four times the baseline amount of human transcriptions. In contrast, standard acoustic data modifications (such as altering pitch or adding noise) and computationally intensive continued pre-training yielded minimal or no meaningful gains.

These results demonstrate that organizations working on language preservation can significantly improve speech recognition systems without incurring the massive labor costs of manual transcription. Self-training leverages readily available unannotated speech recordings, while text-to-speech offers an effective shortcut whenever basic voice synthesizers and written text exist. This framework lowers the technological and financial barriers to building working speech tools for underrepresented communities.

For practical implementation, project teams should prioritize collecting unannotated audio for self-training or utilizing available text-to-speech systems over costly continued pre-training. However, decision-makers should note that increasing synthetic speech volume showed diminishing returns, likely due to single-speaker acoustic bias, and that acoustic performance remains lower than systems for high-resource languages. Future work should focus on multi-speaker synthetic generation, voice conversion methods, and testing broader sociolinguistic variations across larger speaker pools.

Bartelds et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation

Abstract

The performance of automatic speech recognition (ASR) systems has advanced substantially in recent years, particularly for languages for which a large amount of transcribed speech is available. Unfortunately, for low-resource languages, such as minority languages, regional languages or dialects, ASR performance generally remains much lower. In this study, we investigate whether data augmentation techniques could help improve low-resource ASR performance, focusing on four typologically diverse minority languages or language variants (West Germanic: Gronings, West-Frisian; Malayo-Polynesian: Besemah, Nasal). For all four languages, we examine the use of self-training, where an ASR system trained with the available human-transcribed data is used to generate transcriptions, which are then combined with the original data to train a new ASR system. For Gronings, for which there was a pre-existing text-to-speech (TTS) system available, we also examined the use of TTS to generate ASR training data from text-only sources. We find that using a self-training approach consistently yields improved performance (a relative WER reduction up to 20.5% compared to using an ASR system trained on 24 minutes of manually transcribed speech). The performance gain from TTS augmentation for Gronings was even stronger (up to 25.5% relative reduction in WER compared to a system based on 24 minutes of manually transcribed speech). In sum, our results show the benefit of using self-training or (if possible) TTS-generated data as an efficient solution to overcome the limitations of data availability for resource-scarce languages in order to improve ASR performance.

Table of Contents

  • 1 Introduction
  • 2 Data
  • 2.1 Gronings and West-Frisian
  • 2.2 Besemah and Nasal
  • 3 Methods
  • 4 Experimental Setup
  • 4.1 Additional Generated Training Data
  • 4.2 Synthetic speech
  • 5 Results
  • 5.1 Further Pre-Training
  • 5.2 Additional Generated Training Data
  • 5.3 Out-of-domain results
  • 6 Discussion and Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Results on Development Data

Knowls

  1. Knowl 1 — Self-training improves low-resource ASR across four language varieties

    data/table

    The test-set WERs below compare XLS-R models trained with different amounts of manually transcribed speech against student models additionally trained on speech automatically transcribed by a teacher model. Each language had 192 minutes of manually transcribed training speech available; the augmented conditions combine the indicated manually labeled minutes with the remaining portion of that 192-minute training set labeled by self-training. Lower WER is better. The test-set bar charts on page 6 show that self-training reduced WER in nearly every matched condition, with the largest gains generally occurring when the teacher had only 24 minutes of manual transcripts.

    Language 24 min 48 min 96 min 192 min 24+168 ST 48+144 ST 96+96 ST
    Gronings 0.332 0.252 0.202 0.155 0.286 0.230 0.183
    West-Frisian 0.457 0.382 0.307 0.261 0.428 0.352 0.289
    Besemah 0.517 0.423 0.359 0.316 0.471 0.398 0.359
    Nasal 0.591 0.509 0.453 0.413 0.552 0.485 0.437

    At 24 minutes of manual supervision, self-training lowered WER by 13.9% relative for Gronings, 6.3% for West-Frisian, 8.9% for Besemah, and 6.6% for Nasal. The 96-minute Besemah self-training condition tied rather than improved on its manually supervised counterpart. Increasing manual training data from 24 to 192 minutes also reduced WER for every language, by 30.1%–53.3% relative.

  2. Knowl 2 — Synthetic Gronings speech yields large gains, including on out-of-domain speech

    data/table

    For Gronings, the researchers supplemented 24 minutes of manually transcribed speech with synthetic speech generated from text by an existing FastSpeech 2 TTS system. The TTS system had been trained on about two hours of read speech from one female speaker of the Hogelandsters variety. The added text came from recordings corresponding to the remaining Gronings training material. WERs below are for the regular test set and a separate 19-minute out-of-domain test set containing three speakers not used in training. The page-8 paired bar charts compare these evaluation conditions.

    Training data Regular-test WER Out-of-domain WER
    24 min manual 0.332 0.431
    24 min manual + 168 min TTS 0.204 0.321
    24 min manual + 336 min TTS 0.209 0.346
    24 min manual + 672 min TTS 0.198 0.335

    Adding 168 minutes of synthetic speech reduced WER by 38.6% relative on the regular test set and by 25.5% on the out-of-domain test set, compared with the 24-minute manual-only system. Increasing the amount of TTS speech beyond 168 minutes did not yield a consistent additional benefit. The regular test set included a speaker whose speech had been used to train the TTS system, but the improvement persisted on the out-of-domain evaluation.

  3. Knowl 3 — Teacher–student self-training uses unlabeled audio to expand ASR supervision

    model/method

    The self-training procedure starts with a teacher ASR model fine-tuned on a small subset of manually transcribed speech. The teacher transcribes the remaining audio in the available training set, and the resulting pseudo-labeled speech is combined with the original human-labeled examples to train a new student ASR model. For example, a teacher trained on 24 minutes labels the remaining 168 minutes of a 192-minute training set; the student is then trained on the 24 minutes of correct transcripts plus those 168 minutes of automatic transcripts. The same procedure was evaluated with 48 or 96 minutes of initial manual supervision. Decoding used no external language model, both because text resources were limited and to keep the language comparisons consistent.

  4. Knowl 4 — ASR experiments fine-tune multilingual XLS-R under a controlled low-resource setup

    experimental setup

    The ASR system was the 317-million-parameter multilingual XLS-R model, pretrained on approximately 436,000 hours of speech in 128 languages. Its speech encoder and Transformer representations were fine-tuned with a character-level projection and connectionist temporal classification. For fine-tuning, the authors used training subsets of 24, 48, 96, or 192 minutes per language; the 192-minute set was divided into 80% training, 10% development, and 10% test from four hours of manually transcribed speech. Development and test sets were each about 24 minutes. The learning rate was selected from 5×10−45\times10^{-4}, 10−410^{-4}, 5×10−55\times10^{-5}, and 10−510^{-5} using development performance; batch size and gradient accumulation were adjusted for a single 40-GB NVIDIA A100 GPU. The reported test model was the checkpoint with the lowest development WER. The total study used about 390 A100 GPU hours, including fine-tuning and continued-pretraining experiments.

  5. Knowl 5 — Repeated Gronings self-training helps, but added-data gains diminish

    empirical result

    The authors tested iterative self-training for Gronings starting from 24 minutes of manual transcripts. After the initial teacher labeled 168 minutes, each student was used to label additional speech before training the next student; the later additions were 168 minutes and then 336 minutes. The resulting students therefore used 168, 336, or 672 minutes of pseudo-labeled speech in addition to the original 24 minutes. The regular and out-of-domain WERs were:

    Training data Regular-test WER Out-of-domain WER
    24 min manual only 0.332 0.431
    24 min manual + 168 min self-trained 0.286 0.383
    24 min manual + 336 min self-trained 0.281 0.375
    24 min manual + 672 min self-trained 0.264 0.362

    The largest self-training condition reduced regular-test WER by 20.5% relative to the 24-minute manual-only system and out-of-domain WER by 16.0%. Improvements became smaller as more pseudo-labeled material was added; the regular-test performance approached that of a system trained on 48 minutes of manually transcribed speech.

  6. Knowl 6 — The four-language evaluation draws on small documentation and conversational corpora

    data/table

    The study evaluates Gronings and West-Frisian, two West Germanic varieties, and Besemah and Nasal, two Austronesian languages spoken in southern Sumatra. To make training-data amounts comparable, the authors capped manually transcribed data at four hours per language and used the same 80/10/10 split. The table summarizes the principal corpus sources and statistics; the page-2–3 corpus descriptions report these language-specific differences.

    Language Corpus and speech type Utterances Speakers
    Gronings Read speech from three regional varieties; four hours selected from nearly 14 hours of transcribed documentation data 2,130 3
    West-Frisian Radio and television speech selected from the FAME! corpus 4,919 277
    Besemah Informal conversational fieldwork speech 7,835 46
    Nasal Informal conversational fieldwork speech 7,672 40

    The Besemah and Nasal source collections each contained about 45 hours of recorded informal conversation, but only four hours per language were manually transcribed for this study. An additional 19 minutes of Gronings from three other speakers was reserved for out-of-domain testing. The main Gronings data were read aloud by three speakers across Hogelandsters, Oldambtsters, and Westerkwartiers; the West-Frisian corpus comprised broadcast speech from Dutch–Frisian bilinguals.

  7. Knowl 7 — Continued pretraining on limited Gronings audio adds little beyond self-training

    empirical result

    For Gronings, the authors continued pretraining XLS-R on target-language audio for 100,000 steps at a learning rate of 10−510^{-5}, excluding test utterances, before ASR fine-tuning. The test-set WERs below compare the original model with the continued-pretrained (CPT) model under manual-only and self-training conditions.

    Fine-tuning data Original XLS-R WER Continued-pretrained WER
    24 min manual 0.332 0.301
    48 min manual 0.252 0.252
    96 min manual 0.202 0.193
    192 min manual 0.155 0.144
    24 min + 168 min self-trained 0.286 0.282
    48 min + 144 min self-trained 0.230 0.226
    96 min + 96 min self-trained 0.183 0.183

    Continued pretraining improved the 24-minute manual-only result by up to 9.3% relative, but improved self-training conditions by at most 1.7%. Since continued pretraining took about 70 GPU hours and yielded smaller gains than data augmentation, the authors judged it not cost-effective in this setting.

  8. Knowl 8 — Evaluation leaves uncertainty about variance and speaker-related generalization

    limitation

    The experiments used one training run per condition, so the reported WERs have no run-to-run uncertainty estimates or error bars. Each language was evaluated using a single main test set, and the authors note that a different test sample could have changed the measured results. Speaker overlap between train, development, and test splits was allowed because speaker counts were limited. The study also did not measure how sociolinguistic differences or non-linguistic factors such as microphone variation affect recognition; this is especially relevant to Gronings, whose main corpus had only three speakers. Finally, the benefit of augmentation appeared smaller when more manually transcribed speech was available, so the results do not establish that augmentation is always beneficial.

  9. Knowl 9 — Standard signal perturbations did not improve recognition in preliminary trials

    empirical result

    Preliminary experiments tested adding noise to the speech signal, raising or lowering speaker pitch, and simulating far-field speech as training-data augmentation. These techniques did not improve ASR performance in the authors’ experiments, so they were excluded from the main comparison. The paper does not report quantitative WERs for these preliminary trials.

Coverage note — No substantial contributed material was deliberately omitted; the preliminary signal-augmentation result is included, while background and related work are excluded.

References

  1. 1.Alëna Aksënova, Zhehuai Chen, Chung-Cheng Chiu, Daan van Esch, Pavel Golik, Wei Han, Levi King, Bhuvana Ramabhadran, Andrew Rosenberg, Suzan Schwartz, and Gary Wang. 2022. Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data.
  2. 2.Matthew Baas and Herman Kamper. 2022. Voice Conversion Can Improve ASR in Very Low-Resource Settings. In Proc. Interspeech 2022, pages 3513–3517.
  3. 3.Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.
  4. 4.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. In Advances in Neural Information Processing Systems, volume 33, pages 12449–12460. Curran Associates, Inc.
  5. 5.Martijn Bartelds, Wietse de Vries, Faraz Sanal, Caitlin Richter, Mark Liberman, and Martijn Wieling. 2022. Neural representations for modeling variation in speech. Journal of Phonetics, 92:101137.
  6. 6.Martijn Bartelds and Martijn Wieling. 2022. Quantifying language variation acoustically with few resources. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3735–3741, Seattle, United States. Association for Computational Linguistics.
  7. 7.Dan Berrebbi, Ronan Collobert, Samy Bengio, Navdeep Jaitly, and Tatiana Likhomanenko. 2022. Continuous Pseudo-Labeling from the Start.
  8. 8.Rolando Coto-Solano, Sally Akevai Nicholas, Samiha Datta, Victoria Quint, Piripi Wills, Emma Ngakuravaru Powell, Liam Koka’ua, Syed Tanveer, and Isaac Feldman. 2022. Development of automatic speech recognition for the documentation of Cook Islands Maori. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 3872–3882, Marseille, France. European Language Resources Association.
  9. 9.Chenpeng Du and Kai Yu. 2020. Speaker Augmentation for Low Resource Speech Recognition. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7719–7723.
  10. 10.Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. 2013. An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks.
  11. 11.Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Machine Learning, Proceedings of the Twenty-Third International Conference (ICML 2006), Pittsburgh, Pennsylvania, USA, June 25-29, 2006, volume 148 of ACM International Conference Proceeding Series, pages 369–376. ACM.
  12. 12.Séverine Guillaume, Guillaume Wisniewski, Cécile Macaire, Guillaume Jacques, Alexis Michaud, Benjamin Galliot, Maximin Coavoux, Solange Rossato, Minh-Châu Nguyên, and Maxime Fily. 2022. Fine-tuning pre-trained models for automatic speech recognition, experiments on a fieldwork corpus of japhug (trans-himalayan family). In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 170–178, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Jacob Kahn, Ann Lee, and Awni Hannun. 2020. Self-Training for End-to-End Speech Recognition. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7084–7088.
  14. 14.Sameer Khurana, Antoine Laurent, and James Glass. 2022. Magic Dust for Cross-Lingual Adaptation of Monolingual Wav2vec-2.0. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6647–6651.
  15. 15.Loren Lugosch, Tatiana Likhomanenko, Gabriel Synnaeve, and Ronan Collobert. 2022. Pseudo-Labeling for Massively Multilingual Speech Recognition. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7687–7691.
  16. 16.Graham Neubig and Junjie Hu. 2018. Rapid adaptation of neural machine translation to new languages. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 875–880, Brussels, Belgium. Association for Computational Linguistics.
  17. 17.Karol Nowakowski, Michal Ptaszynski, Kyoko Murasaki, and Jagna Nieuwazny. 2023. Adapting multilingual speech representation model for a new, underresourced language through multilingual fine-tuning and continued pretraining. Information Processing & Management, 60(2):103148.
  18. 18.Georgios Paraskevopoulos, Theodoros Kouzelis, Georgios Rouvalis, Athanasios Katsamanis, Vassilis Katsouros, and Alexandros Potamianos. 2023. Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems A case study for Modern Greek.
  19. 19.Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.
  20. 20.Nathaniel Robinson, Perez Ogayo, Swetha Gangu, David R. Mortensen, and Shinji Watanabe. 2022. When Is TTS Augmentation Through a Pivot Language Useful?
  21. 21.Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu. 2019. Speech Recognition with Augmented Synthesized Speech. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 996–1002.
  22. 22.Nick Rossenbach, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2020a. Generating Synthetic Audio Data for Attention-Based Speech Recognition Systems. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7069–7073.
  23. 23.Nick Rossenbach, Albert Zeyer, Ralf Schlüter, and Hermann Ney. 2020b. Generating synthetic audio data for attention-based speech recognition systems. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7069–7073.
  24. 24.Nay San, Martijn Bartelds, Blaine Billings, Ella de Falco, Hendi Feriza, Johan Safri, Wawan Sahrozi, Ben Foley, Bradley McDonnell, and Dan Jurafsky. 2023. Leveraging supplementary text data to kick-start automatic speech recognition system development with limited transcriptions. In Proceedings of the Sixth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 1–6, Remote. Association for Computational Linguistics.
  25. 25.Nay San, Martijn Bartelds, Mitchell Browne, Lily Clifford, Fiona Gibson, John Mansfield, David Nash, Jane Simpson, Myfany Turpin, Maria Vollmer, Sasha Wilmoth, and Dan Jurafsky. 2021. Leveraging Pre-Trained Representations to Improve Access to Untranscribed Speech from Endangered Languages. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1094–1101.
  26. 26.Nay San, Martijn Bartelds, Tolulope Ogunremi, Alison Mount, Ruben Thompson, Michael Higgins, Roy Barker, Jane Simpson, and Dan Jurafsky. 2022. Automated speech tools for helping communities process restricted-access corpora for language revival efforts. In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 41–51, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Anuroop Sriram, Michael Auli, and Alexei Baevski. 2022. Wav2Vec-Aug: Improved self-supervised training with limited data.
  28. 28.Xing Wei, Catia Cucchiarini, Roeland van Hout, and Helmer Strik. 2022. Automatic Speech Recognition and Pronunciation Error Detection of Dutch Non-native Speech: cumulating speech resources in a pluricentric language. Speech Communication, 144:1–9.
  29. 29.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  30. 30.Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021. Self-Training and Pre-Training are Complementary for Speech Recognition. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3030–3034.
  31. 31.Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert. 2020. Iterative Pseudo-Labeling for Speech Recognition. In Proc. Interspeech 2020, pages 1006–1010.
  32. 32.Emre Yılmaz, Jelske Dijkstra, Hans Van de Velde, Frederik Kampstra, Jouke Algra, Henk van den Heuvel, and David Van Leeuwen. 2017. Longitudinal Speaker Clustering and Verification Corpus with Code-Switching Frisian-Dutch Speech. In Proc. Interspeech 2017, pages 37–41.
  33. 33.Emre Yılmaz, Henk van den Heuvel, Jelske Dijkstra, Hans Van de Velde, Frederik Kampstra, Jouke Algra, and David Van Leeuwen. 2016. Open Source Speech and Language Resources for Frisian. In Proc. Interspeech 2016, pages 1536–1540.
  34. 34.Zi-Qiang Zhang, Yan Song, Ming-Hui Wu, Xin Fang, and Li-Rong Dai. 2021. XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition.

Citation

MLA
Bartelds, M., et al. “Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 715–29, https://doi.org/10.18653/v1/2023.acl-long.42.
APA
Bartelds, M., San, N., McDonnell, B., Jurafsky, D., & Wieling, M. (2023). Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 715–729. https://doi.org/10.18653/v1/2023.acl-long.42
Chicago
Bartelds, M., N. San, B. McDonnell, D. Jurafsky, and M. Wieling. 2023. “Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 715–29. https://doi.org/10.18653/v1/2023.acl-long.42.
Harvard
Bartelds, M. et al. (2023) “Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 715–729. Available at: https://doi.org/10.18653/v1/2023.acl-long.42.
Vancouver
1. Bartelds M, San N, McDonnell B, Jurafsky D, Wieling M (2023) Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 715–729

BibTeX

@inproceedings{bartelds-etal-2023-making,
    title = "Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation",
    author = "Bartelds, Martijn  and
      San, Nay  and
      McDonnell, Bradley  and
      Jurafsky, Dan  and
      Wieling, Martijn",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.42/",
    doi = "10.18653/v1/2023.acl-long.42",
    pages = "715--729"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/