Attention as a Guide for Simultaneous Speech Translation
Sara PapiMatteo NegriMarco Turchi
Introduces EDATT, an adaptive policy that applies cross-attention patterns during inference to enable offline-trained speech translation models to perform simultaneous translation with significantly lower latency and higher BLEU scores without dedicated streaming retraining.
Real-time, simultaneous speech translation requires delivering accurate translations with minimal delay as a speaker is talking. Existing approaches typically rely on complex, specialized models that are costly to train across different latency targets, or they employ inefficient decision rules that generate multiple candidate translations before emitting words. To address this efficiency and performance bottleneck, the article evaluates a new decision strategy called EDATT (Encoder-Decoder Attention), which allows standard, offline-trained translation models to perform real-time simultaneous translation without requiring specialized retraining.
The evaluated approach uses the internal attention patterns that naturally form between incoming audio features and target text tokens. By analyzing the sum of attention weights on the most recent audio frames, the system dynamically decides whether the incoming speech contains sufficient information to emit the next translated word or if it must pause and wait for more context. Researchers evaluated this policy on the standard MuST-C benchmark for English-to-German and English-to-Spanish translations using an offline-trained Conformer-Transformer model, comparing performance and latency across multiple metrics and hardware environments.
The findings show that EDATT establishes a superior balance between translation quality and real-time responsiveness. In realistic settings that measure actual computational delay, EDATT achieved quality gains of up to 7.0 BLEU points for German and 4.0 BLEU points for Spanish compared to prior state-of-the-art architectures. Against existing offline policies such as Local Agreement, EDATT reduced computational latency by 0.7 to 1.4 seconds while cutting computational overhead by approximately 30%. Analysis confirmed that extracting attention scores averaged across all heads at the fourth decoder layer, focusing on the last two audio frames, delivered the optimal operational balance across both languages.
These results demonstrate that organizations do not need to invest in costly, purpose-built simultaneous translation architectures to achieve high-performance streaming translation. By applying EDATT directly to existing offline translation models, engineering teams can significantly lower training expenses, simplify system deployment, and reduce inference hardware demands without sacrificing translation quality or user-perceived latency.
Decision-makers should consider adopting EDATT-style inference policies to streamline real-time speech translation pipelines. Before broad deployment, technical teams should validate the optimal frame window and layer configurations on specific domain data, particularly when adapting the framework to non-Western languages or alternative speech encoder compressions. The reported evidence offers high confidence for English-to-European language pairs, provided that realistic, computationally aware latency benchmarks guide final system tuning.
- Paper: Information-Transport-based Policy for Simultaneous Translation, Shaolei Zhang et al. (2022). Its information-transport policy establishes the simultaneous-translation decision problem EDATT revisits, making its latency-versus-information trade-off essential context.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). The Transformer’s encoder-decoder attention architecture supplies the model framework whose attention patterns EDATT repurposes for streaming decisions.
- Paper: Neural Machine Translation by Jointly Learning to Align and Translate, Dzmitry Bahdanau et al. (2015). Bahdanau et al. introduce learned attention alignment during translation, the foundational mechanism EDATT later reads to decide when enough source information has arrived.
- Paper: Effective Approaches to Attention-based Neural Machine Translation, Minh-Thang Luong et al. (2015). Luong et al.’s attention formulations clarify how translation models weight source information while generating target words—the signal EDATT uses to govern emission.
No sufficiently relevant recommendations were found.
