WST: Weakly Supervised Transducer for Automatic Speech Recognition

Dongji GaoChenda LiaoChangliang LiuMatthew WiesnerLeibny Paola GarciaDaniel PoveySanjeev KhudanpurJian Wu

article2025Automatic Speech Recognition & Understanding1 citations

Proposes a weakly supervised transducer framework that trains end-to-end speech recognition models on transcripts with up to 70% error rates using a flexible training graph, outperforming existing temporal classification methods without requiring auxiliary models or confidence scoring.

Listen

Modern automatic speech recognition systems rely heavily on transducer architectures, which require massive volumes of carefully annotated audio to achieve high accuracy. However, human transcription is slow and expensive, especially for low-resource languages, forcing developers to utilize cheap, weakly labeled data such as automated captions or uncurated web recordings. These noisy training sources frequently contain errors—including substituted, inserted, or omitted words—which violate standard training assumptions and severely degrade recognition performance.

The article evaluates and demonstrates a Weakly Supervised Transducer designed to train robust speech recognition models directly from flawed text labels. It aims to eliminate the need for auxiliary pre-trained models or external confidence scoring systems while maintaining state-of-the-art transcription accuracy.

The researchers developed a flexible training graph that automatically bypasses or downweights unreliable text segments by routing potential errors through a generic uncertainty token within an optimized graph framework. To maintain manageable computational complexity, the system simplifies decoder history by evaluating only the most recent token. The authors assessed the architecture across controlled benchmarks using the standard LibriSpeech dataset injected with synthetic transcription errors ranging from 10% to 70%, as well as a large-scale industrial dataset containing 10,000 hours of realistic, noisy English speech evaluated across multiple regional accents.

The evaluation yielded several critical findings. First, under clean training conditions, the proposed model matches the performance of standard transducers without penalty. Second, as noise increases, the model demonstrates remarkable resilience; it successfully trains under synthetic error rates as high as 70%, whereas standard connectionist models completely fail to converge. Third, in extreme substitution error scenarios on benchmark data, the model achieved a 13.0% word error rate compared to 21.5% for previous weak-supervision baselines, representing a relative error reduction of nearly 40%. Finally, on the 10,000-hour industrial dataset containing naturally occurring noise, the approach reduced overall recognition errors from 22.65% to 21.71% across various regional and non-native accents, delivering relative improvements of up to 9.71% in specific accented subsets.

These results demonstrate that organizations can significantly lower speech recognition training costs and timelines by safely ingesting lower-quality, uncurated data at scale. Rather than expending substantial resources on manual data curation or multi-stage filtering pipelines, teams can train robust end-to-end models directly from scratch. The framework effectively mitigates deployment risks across diverse acoustic environments and varied speaker accents.

Engineering and product leaders building speech recognition systems should consider replacing conventional transducer loss layers with this weakly supervised framework when dealing with imperfect transcripts. Development teams can implement the technique directly into standard training workflows to exploit large backlogs of uncurated audio data. Before full-scale enterprise rollout, organizations should conduct pilot evaluations on their specific operational data to establish optimal arc penalty settings and confirm that localized dialect performance remains consistent across target demographics.

Confidence in these findings is high given the consistent performance across both synthetic benchmarks and large-scale industrial data. However, minor limitations remain: the framework assumes the number of missing tokens does not exceed the total audio frames, relies on a simplified decoder history approximation, and observed a negligible performance degradation of 0.10% in one accented evaluation subset. Practitioners should review domain-specific data profiles to ensure alignment with these operational assumptions.

No sufficiently relevant recommendations were found.

Cover for WST: Weakly Supervised Transducer for Automatic Speech Recognition

Abstract

The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available.

Table of Contents

  • I Introduction
  • II Preliminaries
  • II-A Transducer
  • II-B Transducer in WFST framework
  • III Method
  • III-A Weakly Supervised Transcript Graph
  • III-B Weakly Supervised Training Graph
  • III-C Decoder History Approximation
  • III-D Modeling ⋆\star token
  • III-E Arc weight (penalty) strategy
  • IV Experimental Setup
  • IV-A Data Preparation
  • IV-B Implementation Details
  • V Results and Analysis
  • V-A LibriSpeech
  • V-B IH-10k
  • VI Relation to Concurrent Work
  • VII Conlusions
  • References

Knowls

  1. Knowl 1 — WST adds wildcard bypass paths to the transducer training graph

    model/method

    Weakly Supervised Transducer (WST) represents a possibly corrupted transcript as a weighted finite-state transducer (WFST), then expands that transcript graph over acoustic time to form the training graph. In a standard transducer graph, blank arcs consume an acoustic frame without emitting a label, while token arcs emit a label without advancing acoustic time. WST adds a special wildcard token ⋆\star and two bypass mechanisms: token bypass arcs move along the transcript axis, allowing a transcript token to be skipped within the same acoustic frame; blank bypass arcs move along the time axis, consuming an acoustic frame while emitting ⋆\star. Token bypass paths can avoid learning from incorrect or extraneous transcript tokens, and blank bypass paths can account for speech content missing from the transcript. The graph therefore represents alternatives for substitution, insertion, and deletion errors, and the transducer objective sums the probabilities of its permitted alignment paths. WST is designed to train from scratch without an external confidence estimator or pretrained ASR model. Its deletion mechanism assumes that the number of missing transcript tokens does not exceed the number TT of acoustic frames.

  2. Knowl 2 — Wildcard probability distributes nonblank probability mass uniformly

    equation

    At acoustic frame tt, WST assigns the wildcard token ⋆\star the average probability of the nonblank output symbols:

    P(⋆t∣x)=1−P(blankt∣x)∣V∣−1.P(\star_t\mid x)=\frac{1-P(\text{blank}_t\mid x)}{|V|-1}.

    Here, xx is the acoustic input, P(blankt∣x)P(\text{blank}_t\mid x) is the model's blank probability at frame tt, and VV is the output-symbol inventory including the blank symbol, so ∣V∣−1|V|-1 is the number of nonblank symbols. Thus the model's nonblank probability mass at each frame is divided equally among those symbols to obtain the wildcard probability.

  3. Knowl 3 — WST approximates branching decoder histories with a stateless predictor

    model/method

    A standard autoregressive transducer predictor conditions its representation at output position uu on the full preceding label history (y1,…,yu−1)(y_1,\ldots,y_{u-1}). WST's wildcard paths make the transcript graph branch, so different label histories can reach the same graph state. To avoid maintaining a distinct predictor state for every history, WST uses a stateless prediction network (SLP) whose representation depends only on the immediately preceding symbol:

    gu=SLP⁡(yu−1).g_u=\operatorname{SLP}(y_{u-1}).

    Here, gug_u is the predictor representation at output position uu, and yu−1y_{u-1} is the preceding symbol, which may be ⋆\star or a vocabulary token. WST approximates the alternative histories reaching a state as sharing a predictor representation rather than carrying separate history-conditioned representations. This permits training with the branching graph without adding history-tracking complexity to the decoder.

  4. Knowl 4 — WST uses constant penalties for both bypass-arc types

    model/method

    WST assigns separate fixed penalties to token bypass arcs and blank bypass arcs. For training epoch ii, the penalties are λi(1)=β1\lambda_i^{(1)}=\beta_1 for token bypass arcs and λi(2)=β2\lambda_i^{(2)}=\beta_2 for blank bypass arcs, where β1\beta_1 and β2\beta_2 are constants that do not change across epochs. This fixed-penalty strategy avoids an epoch-dependent decay schedule.

  5. Knowl 5 — Evaluation data and transducer configurations

    experimental setup

    Experiments use LibriSpeech train-clean-100 with synthetic transcript errors at rates 0.10.1, 0.30.3, 0.50.5, and 0.70.7, in separate substitution-only, insertion-only, deletion-only, and mixed-error conditions. LibriSpeech inputs are 768-dimensional features extracted using wav2vec 2.0 base at a 20 ms stride, and the BPE vocabulary has 200 symbols. The second training corpus, IH-10k, contains 10,000 hours of industrial audio; it uses 80-dimensional Fbank features and a 4,000-symbol BPE vocabulary. A manual assessment of 400 IH-10k utterances estimates a transcript error rate of about 10.0%, comprising 4.0% substitutions, 2.0% insertions, and 4.0% deletions. The transducer uses a 12-layer Conformer encoder, a stateless feed-forward prediction network, and a fully connected joiner followed by softmax. LibriSpeech comparisons include CTC, BTC, OTC, a standard transducer, and WST; the IH-10k comparison is between the standard transducer and WST. LibriSpeech word error rates (WERs) are measured with greedy decoding.

  6. Knowl 6 — LibriSpeech test-clean results under synthetic transcript corruption

    empirical result

    The following are greedy-decoding WERs in percent on LibriSpeech test-clean, after training on train-clean-100 with the indicated synthetic error type. Within each vector, values correspond in order to transcript error rates 0.0,0.1,0.3,0.5,0.70.0, 0.1, 0.3, 0.5, 0.7. A dash denotes failure to converge; N/A denotes that BTC is not applicable to deletion errors.

    • Substitution: CTC [7.8,15.1,20.8,47.7,−][7.8, 15.1, 20.8, 47.7, -]; BTC [7.8,14.7,17.5,19.8,−][7.8, 14.7, 17.5, 19.8, -]; OTC [7.8,8.9,11.0,15.4,21.5][7.8, 8.9, 11.0, 15.4, 21.5]; standard transducer [7.1,9.5,12.0,29.4,−][7.1, 9.5, 12.0, 29.4, -]; WST [7.1,8.3,9.0,10.4,13.0][7.1, 8.3, 9.0, 10.4, 13.0].
    • Insertion: CTC [7.8,18.7,29.8,72.8,−][7.8, 18.7, 29.8, 72.8, -]; BTC [7.8,12.0,12.1,12.1,12.7][7.8, 12.0, 12.1, 12.1, 12.7]; OTC [7.8,7.8,7.9,7.9,8.0][7.8, 7.8, 7.9, 7.9, 8.0]; standard transducer [7.1,8.6,9.5,11.2,14.6][7.1, 8.6, 9.5, 11.2, 14.6]; WST [7.1,7.3,7.5,7.7,7.8][7.1, 7.3, 7.5, 7.7, 7.8].
    • Deletion: CTC [7.8,9.8,17.2,57.7,−][7.8, 9.8, 17.2, 57.7, -]; BTC [7.8,N/A,N/A,N/A,N/A][7.8, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}]; OTC [7.8,8.3,10.3,17.6,22.9][7.8, 8.3, 10.3, 17.6, 22.9]; standard transducer [7.1,9.2,21.7,−,−][7.1, 9.2, 21.7, -, -]; WST [7.1,8.3,10.4,15.8,21.6][7.1, 8.3, 10.4, 15.8, 21.6].
    • Mixed errors: CTC [7.8,10,17.2,−,−][7.8, 10, 17.2, -, -]; BTC [7.8,N/A,N/A,N/A,N/A][7.8, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}]; OTC [7.8,8.5,9.9,13.1,29.4][7.8, 8.5, 9.9, 13.1, 29.4]; standard transducer [7.1,8.7,11.2,27.1,−][7.1, 8.7, 11.2, 27.1, -]; WST [7.1,8.1,8.6,10.2,19.1][7.1, 8.1, 8.6, 10.2, 19.1].

    WST matches the standard transducer's clean-data WER of 7.1% and has the lowest WER among the reported systems at the 70% error level in all four conditions. For example, at 70% substitution error, WST obtains 13.0% WER versus 21.5% for OTC; at 70% insertion error, the corresponding values are 7.8% and 8.0%. The standard transducer remains trainable at 70% insertion error, with 14.6% WER, while CTC fails to converge.

  7. Knowl 7 — LibriSpeech test-other results under synthetic transcript corruption

    empirical result

    The following are greedy-decoding WERs in percent on LibriSpeech test-other, after training on train-clean-100 with the indicated synthetic error type. Within each vector, values correspond in order to transcript error rates 0.0,0.1,0.3,0.5,0.70.0, 0.1, 0.3, 0.5, 0.7. A dash denotes failure to converge; N/A denotes that BTC is not applicable to deletion errors.

    • Substitution: CTC [19.4,29.2,36.5,60.5,−][19.4, 29.2, 36.5, 60.5, -]; BTC [19.4,29.0,33.3,36.5,−][19.4, 29.0, 33.3, 36.5, -]; OTC [19.4,21.2,24.7,32.3,39.1][19.4, 21.2, 24.7, 32.3, 39.1]; standard transducer [17.8,22.7,29.3,40.5,−][17.8, 22.7, 29.3, 40.5, -]; WST [17.8,20.1,21.8,23.6,26.9][17.8, 20.1, 21.8, 23.6, 26.9].
    • Insertion: CTC [19.4,31.0,44.3,−,−][19.4, 31.0, 44.3, -, -]; BTC [19.4,24.0,24.1,24.4,24.6][19.4, 24.0, 24.1, 24.4, 24.6]; OTC [19.8,19.4,19.5,19.6,21.4][19.8, 19.4, 19.5, 19.6, 21.4]; standard transducer [17.8,20.6,22.1,25.1,28.5][17.8, 20.6, 22.1, 25.1, 28.5]; WST [17.8,18.1,18.2,18.9,19.3][17.8, 18.1, 18.2, 18.9, 19.3].
    • Deletion: CTC [19.4,21.5,29.4,−,−][19.4, 21.5, 29.4, -, -]; BTC [19.4,N/A,N/A,N/A,N/A][19.4, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}]; OTC [19.4,20.7,23.7,33.2,41.8][19.4, 20.7, 23.7, 33.2, 41.8]; standard transducer [17.8,20.5,31.5,−,−][17.8, 20.5, 31.5, -, -]; WST [17.8,20.3,24.0,28.7,39.5][17.8, 20.3, 24.0, 28.7, 39.5].
    • Mixed errors: CTC [19.4,21.5,28.0,−,−][19.4, 21.5, 28.0, -, -]; BTC [19.4,N/A,N/A,N/A,N/A][19.4, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}, \mathrm{N/A}]; OTC [19.4,20.4,22.3,26.4,44.6][19.4, 20.4, 22.3, 26.4, 44.6]; standard transducer [17.8,21.0,23.7,37.1,−][17.8, 21.0, 23.7, 37.1, -]; WST [17.8,19.2,20.2,24.2,30.4][17.8, 19.2, 20.2, 24.2, 30.4].

    At zero synthetic error, WST matches the standard transducer at 17.8% WER. At 70% error, WST has the lowest reported WER in each condition: 26.9% for substitutions, 19.3% for insertions, 39.5% for deletions, and 30.4% for mixed errors. In the 70% insertion condition, the standard transducer reaches 28.5% WER while CTC fails to converge.

  8. Knowl 8 — WST improves WER on most IH-10k accent subsets

    empirical result

    On an internal benchmark spanning six English-accent subsets, models trained on the 10,000-hour IH-10k corpus were evaluated by WER (percent). WST lowers WER relative to the standard transducer on five of the six subsets: en-Accent-1, 19.78 to 18.93 (4.32% relative improvement); en-Accent-2, 23.98 to 22.75 (5.12%); en-Accent-3, 16.35 to 14.76 (9.71%); en-Accent-4, 28.78 to 27.27 (5.25%); and en-Accent-6, 22.79 to 21.87 (4.03%). On en-Accent-5, WER changes from 24.48 to 24.50, a 0.10% relative performance loss. Across the full benchmark, WER decreases from 22.65% for the standard transducer to 21.71% for WST, a 4.15% relative improvement. The accent groups include non-native and regionally varied English, and the IH-10k training transcripts were estimated to contain about 10% errors from manual annotation of 400 utterances.

Coverage note — No substantial contributed material is omitted; standard transducer background and the related/concurrent-work discussion are excluded as non-contribution material.

References

  1. 1.A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML, 2006.
  2. 2.W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP. IEEE, 2016.
  3. 3.A. Graves, “Sequence transduction with recurrent neural networks,” in ICML, 2012.
  4. 4.S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid ctc/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing, 2017.
  5. 5.J. Li et al., “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing, vol. 11, no. 1, 2022.
  6. 6.J. Li, Y. Wu, Y. Gaur, C. Wang, R. Zhao, and S. Liu, “On the comparison of popular end-to-end models for large scale speech recognition,” in INTERSPEECH, 2020.
  7. 7.L. Lu, C. Liu, J. Li, and Y. Gong, “Exploring transformers for large-scale speech recognition,” in INTERSPEECH, 2020.
  8. 8.Y. Wang, Y. Shi, F. Zhang, C. Wu, J. Chan, C.-F. Yeh, and A. Xiao, “Transformer in action: A comparative study of transformer-based acoustic models for large scale speech recognition applications,” in ICASSP. IEEE, 2020.
  9. 9.A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML. PMLR, 2023.
  10. 10.V. Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi et al., “Scaling speech technology to 1,000+ languages,” Journal of Machine Learning Research, 2024.
  11. 11.Y. Peng, J. Tian, B. Yan, D. Berrebbi, X. Chang, X. Li, J. Shi, S. Arora, W. Chen, R. Sharma et al., “Reproducing whisper-style training using an open-source toolkit and publicly available data,” in ASRU. IEEE, 2023.
  12. 12.X. Cai, J. Yuan, Y. Bian, G. Xun, J. Huang, and K. Church, “W-ctc: a connectionist temporal classification loss with wild cards,” in ICLR, 2022.
  13. 13.V. Pratap, A. Hannun, G. Synnaeve, and R. Collobert, “Star temporal classification: Sequence classification with partially labeled data,” in NeurIPS, 2022.
  14. 14.D. Gao, M. Wiesner, H. Xu, L. P. Garcia, D. Povey, and S. Khudanpur, “Bypass temporal classification: Weakly supervised automatic speech recognition with imperfect transcripts,” in INTERSPEECH, 2023.
  15. 15.D. Gao, H. Xu, D. Raj, L. P. G. Perera, D. Povey, and S. Khudanpur, “Learning from flawed data: Weakly supervised automatic speech recognition,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023.
  16. 16.A. Laptev, V. Bataev, I. Gitman, and B. Ginsburg, “Powerful and extensible wfst framework for rnn-transducer losses,” in ICASSP, 2023.
  17. 17.G. Keren, W. Zhou, and O. Kalinli, “Token-weighted rnn-t for learning from flawed data,” in SLT, 2024.
  18. 18.H. Zhu, D. Gao, G. Cheng, D. Povey, P. Zhang, and Y. Yan, “Alternative pseudo-labeling for semi-supervised automatic speech recognition,” TASLP, 2023.
  19. 19.D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi.” in Interspeech, 2016.
  20. 20.M. Ghodsi, X. Liu, J. Apfel, R. Cabrera, and E. Weinstein, “Rnn-transducer with stateless prediction network,” in ICASSP, 2020.
  21. 21.V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP. IEEE, 2015.
  22. 22.A. Baevski, Y. Zhou, A. Mohamed et al., “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS, 2020.
  23. 23.A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al., “Conformer: Convolution-augmented transformer for speech recognition,” INTERSPEECH, 2020.

Citation

MLA
Gao, D., et al. “WST: Weakly Supervised Transducer for Automatic Speech Recognition”. arXiv, 2025, http://arxiv.org/abs/2511.04035v1.
APA
Gao, D., Liao, C., Liu, C., Wiesner, M., Garcia, L. P., Povey, D., Khudanpur, S., & Wu, J. (2025). WST: Weakly Supervised Transducer for Automatic Speech Recognition. arXiv. http://arxiv.org/abs/2511.04035v1
Chicago
Gao, D., C. Liao, C. Liu, et al. 2025. “WST: Weakly Supervised Transducer for Automatic Speech Recognition”. arXiv. http://arxiv.org/abs/2511.04035v1.
Harvard
Gao, D. et al. (2025) “WST: Weakly Supervised Transducer for Automatic Speech Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2511.04035v1.
Vancouver
1. Gao D, Liao C, Liu C, Wiesner M, Garcia LP, Povey D, Khudanpur S, Wu J (2025) WST: Weakly Supervised Transducer for Automatic Speech Recognition. arXiv

BibTeX

@article{gao2025wst,
  title = {WST: Weakly Supervised Transducer for Automatic Speech Recognition},
  author = {Gao, Dongji and Liao, Chenda and Liu, Changliang and Wiesner, Matthew and Garcia, Leibny Paola and Povey, Daniel and Khudanpur, Sanjeev and Wu, Jian},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2511.04035v1},
  eprint = {2511.04035}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/