WST: Weakly Supervised Transducer for Automatic Speech Recognition
Dongji GaoChenda LiaoChangliang LiuMatthew WiesnerLeibny Paola GarciaDaniel PoveySanjeev KhudanpurJian Wu
Proposes a weakly supervised transducer framework that trains end-to-end speech recognition models on transcripts with up to 70% error rates using a flexible training graph, outperforming existing temporal classification methods without requiring auxiliary models or confidence scoring.
Modern automatic speech recognition systems rely heavily on transducer architectures, which require massive volumes of carefully annotated audio to achieve high accuracy. However, human transcription is slow and expensive, especially for low-resource languages, forcing developers to utilize cheap, weakly labeled data such as automated captions or uncurated web recordings. These noisy training sources frequently contain errors—including substituted, inserted, or omitted words—which violate standard training assumptions and severely degrade recognition performance.
The article evaluates and demonstrates a Weakly Supervised Transducer designed to train robust speech recognition models directly from flawed text labels. It aims to eliminate the need for auxiliary pre-trained models or external confidence scoring systems while maintaining state-of-the-art transcription accuracy.
The researchers developed a flexible training graph that automatically bypasses or downweights unreliable text segments by routing potential errors through a generic uncertainty token within an optimized graph framework. To maintain manageable computational complexity, the system simplifies decoder history by evaluating only the most recent token. The authors assessed the architecture across controlled benchmarks using the standard LibriSpeech dataset injected with synthetic transcription errors ranging from 10% to 70%, as well as a large-scale industrial dataset containing 10,000 hours of realistic, noisy English speech evaluated across multiple regional accents.
The evaluation yielded several critical findings. First, under clean training conditions, the proposed model matches the performance of standard transducers without penalty. Second, as noise increases, the model demonstrates remarkable resilience; it successfully trains under synthetic error rates as high as 70%, whereas standard connectionist models completely fail to converge. Third, in extreme substitution error scenarios on benchmark data, the model achieved a 13.0% word error rate compared to 21.5% for previous weak-supervision baselines, representing a relative error reduction of nearly 40%. Finally, on the 10,000-hour industrial dataset containing naturally occurring noise, the approach reduced overall recognition errors from 22.65% to 21.71% across various regional and non-native accents, delivering relative improvements of up to 9.71% in specific accented subsets.
These results demonstrate that organizations can significantly lower speech recognition training costs and timelines by safely ingesting lower-quality, uncurated data at scale. Rather than expending substantial resources on manual data curation or multi-stage filtering pipelines, teams can train robust end-to-end models directly from scratch. The framework effectively mitigates deployment risks across diverse acoustic environments and varied speaker accents.
Engineering and product leaders building speech recognition systems should consider replacing conventional transducer loss layers with this weakly supervised framework when dealing with imperfect transcripts. Development teams can implement the technique directly into standard training workflows to exploit large backlogs of uncurated audio data. Before full-scale enterprise rollout, organizations should conduct pilot evaluations on their specific operational data to establish optimal arc penalty settings and confirm that localized dialect performance remains consistent across target demographics.
Confidence in these findings is high given the consistent performance across both synthetic benchmarks and large-scale industrial data. However, minor limitations remain: the framework assumes the number of missing tokens does not exceed the total audio frames, relies on a simplified decoder history approximation, and observed a negligible performance degradation of 0.10% in one accented evaluation subset. Practitioners should review domain-specific data profiles to ensure alignment with these operational assumptions.
- Paper: Sequence Transduction with Recurrent Neural Networks, Alex Graves (2012). Graves establishes the RNN-Transducer’s joint input-output sequence model, the central architecture that WST modifies to tolerate noisy transcripts.
- Paper: Speech Recognition with Deep Recurrent Neural Networks, Alex Graves et al. (2013). This work develops the deep RNN-Transducer and CTC foundations whose alignment and prediction mechanisms WST adapts for weak supervision.
- Paper: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin, Dario Amodei et al. (2016). Deep Speech 2 clarifies the CTC-based ASR alternative that WST’s comparisons build on and seek to surpass under transcript noise.
- Paper: Robust Speech Recognition via Large-Scale Weak Supervision, Alec Radford et al. (2023). Whisper demonstrates large-scale ASR training with imperfect web transcripts, motivating the weak-supervision setting that WST addresses with a transducer-specific training graph.
No sufficiently relevant recommendations were found.
