Attention as a Guide for Simultaneous Speech Translation

Sara PapiMatteo NegriMarco Turchi

article2023ACL42 citations

Introduces EDATT, an adaptive policy that applies cross-attention patterns during inference to enable offline-trained speech translation models to perform simultaneous translation with significantly lower latency and higher BLEU scores without dedicated streaming retraining.

Listen

Real-time, simultaneous speech translation requires delivering accurate translations with minimal delay as a speaker is talking. Existing approaches typically rely on complex, specialized models that are costly to train across different latency targets, or they employ inefficient decision rules that generate multiple candidate translations before emitting words. To address this efficiency and performance bottleneck, the article evaluates a new decision strategy called EDATT (Encoder-Decoder Attention), which allows standard, offline-trained translation models to perform real-time simultaneous translation without requiring specialized retraining.

The evaluated approach uses the internal attention patterns that naturally form between incoming audio features and target text tokens. By analyzing the sum of attention weights on the most recent audio frames, the system dynamically decides whether the incoming speech contains sufficient information to emit the next translated word or if it must pause and wait for more context. Researchers evaluated this policy on the standard MuST-C benchmark for English-to-German and English-to-Spanish translations using an offline-trained Conformer-Transformer model, comparing performance and latency across multiple metrics and hardware environments.

The findings show that EDATT establishes a superior balance between translation quality and real-time responsiveness. In realistic settings that measure actual computational delay, EDATT achieved quality gains of up to 7.0 BLEU points for German and 4.0 BLEU points for Spanish compared to prior state-of-the-art architectures. Against existing offline policies such as Local Agreement, EDATT reduced computational latency by 0.7 to 1.4 seconds while cutting computational overhead by approximately 30%. Analysis confirmed that extracting attention scores averaged across all heads at the fourth decoder layer, focusing on the last two audio frames, delivered the optimal operational balance across both languages.

These results demonstrate that organizations do not need to invest in costly, purpose-built simultaneous translation architectures to achieve high-performance streaming translation. By applying EDATT directly to existing offline translation models, engineering teams can significantly lower training expenses, simplify system deployment, and reduce inference hardware demands without sacrificing translation quality or user-perceived latency.

Decision-makers should consider adopting EDATT-style inference policies to streamline real-time speech translation pipelines. Before broad deployment, technical teams should validate the optimal frame window and layer configurations on specific domain data, particularly when adapting the framework to non-Western languages or alternative speech encoder compressions. The reported evidence offers high confidence for English-to-European language pairs, provided that realistic, computationally aware latency benchmarks guide final system tuning.

arXiv: 2212.07850

No sufficiently relevant recommendations were found.

Cover for Attention as a Guide for Simultaneous Speech Translation

Abstract

In simultaneous speech translation (SimulST), effective policies that determine when to write partial translations are crucial to reach high output quality with low latency. Towards this objective, we propose EDATT (Encoder-Decoder Attention), an adaptive policy that exploits the attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference. EDATT exploits the attention scores modeling the audio-translation relation to decide whether to emit a partial hypothesis or wait for more audio input. This is done under the assumption that, if attention is focused towards the most recently received speech segments, the information they provide can be insufficient to generate the hypothesis (indicating that the system has to wait for additional audio input). Results on en→{de, es} show that EDATT yields better results compared to the SimulST state of the art, with gains respectively up to 7 and 4 BLEU points for the two languages, and with a reduction in computational-aware latency up to 1.4s and 0.7s compared to existing SimulST policies applied to offline-trained models.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 EDAtt policy
  • 4 Experimental Settings
  • 4.1 Data
  • 4.2 Architecture and Training Setup
  • 4.3 Inference and Evaluation
  • 4.4 Terms of Comparison
  • 5 Attention Analysis
  • 6 Results
  • 6.1 Comparison with Other Approaches
  • 6.2 Effects of Accelerated Hardware
  • 7 Related Works
  • 8 Conclusions
  • References
  • A Training Settings
  • B Data Statistics
  • C Main Results with Different Latency Metrics
  • D Numeric Values for Main Results

Knowls

  1. Knowl 1 — EDATT uses encoder–decoder attention to decide when to emit

    algorithm

    EDATT (Encoder-Decoder Attention) is an adaptive simultaneous speech translation policy that uses the attention weights of an offline-trained speech translation model, without policy-specific retraining or architectural changes. For a partial audio input ending at encoder position tt, let yjy_j be the next candidate target token, and let Ai,jA_{i,j} be the encoder–decoder attention weight from that token to source position ii. EDATT emits yjy_j when the attention mass over the final source positions is below a threshold; otherwise it waits for more audio. In the reported configuration, attention is averaged across all eight heads in decoder layer 4, the speech arrives in 800 ms segments, and the number of final positions is set to λ=2\lambda=2.

    ∑i=t−λtAi,j(Xt,Yj−1)<α,α∈(0,1)\sum_{i=t-\lambda}^{t} A_{i,j}(X_t,Y_{j-1}) < \alpha, \qquad \alpha\in(0,1)

    Here, XtX_t is the audio prefix available at the current step, Yj−1Y_{j-1} is the sequence of previously emitted target tokens, tt is the last available encoder position, and α\alpha is the emission threshold. The inequality is the paper’s stated decision rule: a candidate is emitted if its attention is not sufficiently concentrated at the end of the available audio. The tested threshold values were 0.6,0.4,0.2,0.1,0.05,0.6, 0.4, 0.2, 0.1, 0.05, and 0.030.03; lower thresholds generally increase latency.

    Input: Audio stream, offline-trained speech translation model, threshold α\alpha
    For each newly received 800 ms audio segment:
        Update the model input to include the available audio prefix XtX_t
        Compute encoder-decoder attention, averaging heads in decoder layer 4
        While a next target-token candidate yjy_j is available:
            If ∑i=t−λtAi,j(Xt,Yj−1)<α\sum_{i=t-\lambda}^{t} A_{i,j}(X_t,Y_{j-1}) < \alpha:
                Emit yjy_j and continue to the next candidate
            Else:
                Stop emitting and wait for more audio
    Output: The incrementally emitted target translation
  2. Knowl 2 — EDATT outperforms Local Agreement and Wait-k on offline models

    empirical result

    On MuST-C English-to-German and English-to-Spanish tst-COMMON, EDATT was compared with Local Agreement (LA) and Wait-k when these policies were applied to the same offline-trained speech translation models. Translation quality was measured with BLEU; computational-aware average lagging (AL_CA) measures elapsed latency including computation, unlike idealized average lagging (AL). Across almost all latency points, EDATT improved on LA: the paper reports AL_CA reductions of up to 1.4 seconds for German and 0.7 seconds for Spanish, with BLEU gains of up to 2.0 and 3.0 points, respectively. EDATT’s computational overhead averaged 0.9 seconds across the languages, compared with 1.3 seconds for LA, a reported 30% reduction. Against Wait-k, EDATT gained 1.0–2.5 BLEU for German and 1.0–3.0 for Spanish across the reported latency measures. EDATT also began producing output at about 1.0 second of latency, compared with about 1.5 seconds for Wait-k.

  3. Knowl 3 — EDATT reveals source–target structure in speech translation attention

    empirical result

    Inspection of encoder–decoder attention from offline-trained models showed a recurring pattern in MuST-C English-to-German and English-to-Spanish: attention weights concentrated on the final input frame, largely regardless of audio length. In the visualization analysis, removing that final frame from the attention matrix and renormalizing the remaining weights exposed a clearer pseudo-diagonal pattern, linking source audio positions with target translation positions. This observation supports the premise behind EDATT—that attention can indicate whether a target-token candidate depends on the most recently received audio. The final-frame removal was used to clarify the visualization; it was not part of EDATT’s emission rule.

  4. Knowl 4 — Experimental data, offline model, and evaluation conditions

    experimental setup

    The experiments used MuST-C English-to-German and English-to-Spanish speech translation data, with sequence-level knowledge distillation: translated training transcripts were used alongside gold references, and the reported training counts reflect the doubling from this procedure. The effective training counts were 225,277 examples for German and 260,049 for Spanish; dev and tst-COMMON contained 1,423 and 1,422 examples for German, and 1,316 and 1,315 for Spanish. Training samples with audio longer than 30 seconds were discarded.

    The offline speech translation model had 12 Conformer encoder layers and 6 Transformer decoder layers, about 115 million parameters, and 8 attention heads per layer. Its input consisted of 80 audio features computed every 10 ms using a 25 ms window; two stride-2 convolutional layers reduced the sequence length by a factor of four. Simultaneous inference was simulated with SimulEval, using 800 ms audio segments for EDATT, on a single NVIDIA K80 GPU with 12 GB of memory. BLEU measured translation quality; AL and AL_CA measured latency.

  5. Knowl 5 — The selected attention window is two encoder positions

    empirical result

    The authors selected the EDATT window size λ\lambda using the MuST-C dev sets for both language directions. They tested λ∈{2,4,6,8}\lambda\in\{2,4,6,8\}, varying the threshold α\alpha and averaging attention across heads in decoder layer 5 for this window-size analysis. Increasing λ\lambda shifted the quality–latency curves toward higher latency, while values of 6 or 8 provided little quality benefit relative to that latency cost. The choice λ=2\lambda=2 gave the lowest reported latency, about 1.2 seconds of AL in both languages, and was therefore used in subsequent experiments. The authors also report that λ=1\lambda=1 consistently degraded translation quality. This shared setting was observed for these two language directions, not established as universal across languages or model architectures.

  6. Knowl 6 — Decoder layer 4 with head-averaged attention is the practical setting

    empirical result

    With λ=2\lambda=2, the authors compared attention from each of the six decoder layers on the MuST-C dev sets. Layers 1 and 2 performed worse than later layers, and layer 3 generally had lower quality than layers 4–6. Layers 5 and 6 could perform well at higher latency but did not reach the lowest-latency region: layer 6 began around 1.5 seconds of AL, and layer 5 showed a smaller version of this constraint. Layer 4 was consequently chosen for the simultaneous setting. They also compared the eight attention heads in layer 4: no individual head was consistently best across languages and latency points, and many heads could not reach low latency. Averaging across all eight heads gave better overall quality–latency performance than using an individual head.

  7. Knowl 7 — EDATT’s comparison with CAAT depends on the latency measure

    empirical result

    On MuST-C English-to-German and English-to-Spanish tst-COMMON, EDATT was also compared with CAAT, a jointly trained simultaneous translation architecture. Under idealized AL, EDATT had higher BLEU at medium-to-high latency (AL at least 1.2 seconds), with gains up to 5.0 points for German and 2.0 for Spanish; below 1.2 seconds, EDATT’s BLEU was lower by 1.5–4.0 points for German and 1.0–2.5 for Spanish. Under computational-aware AL_CA, the authors report that EDATT’s curves were always to the left of CAAT’s and give gains up to 6.0 BLEU for German and 2.0 for Spanish. The abstract and conclusion state headline gains of up to 7.0 and 4.0 BLEU, respectively, against the state of the art. These results are reported for different summaries or conditions in the paper and should not be treated as a single matched-latency estimate.

  8. Knowl 8 — EDATT retains its advantage with accelerated inference hardware

    empirical result

    The authors reran simultaneous inference on an NVIDIA A40 GPU with 48 GB of memory to examine hardware effects on computational-aware latency. Relative to the K80 results, LA benefited most, reducing latency by about 0.5–1.0 seconds, but still did not reach latency below 2 seconds. Wait-k and CAAT shifted by less than 0.5 seconds, as did EDATT. EDATT remained superior in the reported quality–latency comparison on the A40. The authors note that the A40’s higher cost (reported as 4.1perhourversus4.1 per hour versus 0.9 per hour for the K80) limits the practical appeal of obtaining that acceleration through more expensive hardware.

  9. Knowl 9 — EDATT’s tested scope and window size are limited

    limitation

    The paper analyzes EDATT on speech translation models using CTC compression, which changes the encoder sequence into shorter, more meaningful units. The appropriate value of λ\lambda can therefore depend on the model, and the authors state that it should be selected on a validation set before testing. Experiments covered English as the source language and Western European target languages (German and Spanish); application to other source languages or non-Western European targets was not verified.

Coverage note — Detailed training hyperparameters and supplementary DAL/LAAL metric analyses are omitted because they provide implementation or secondary evaluation detail rather than additional central contributions.

References

  1. 1.Antonios Anastasopoulos, Loïc Barrault, Luisa Bentivogli, Marcely Zanon Boito, Ondřej Bojar, Roldano Cattoni, Anna Currey, Georgiana Dinu, Kevin Duh, Maha Elbayad, Clara Emmanuel, Yannick Estève, Marcello Federico, Christian Federmann, Souhir Gahbiche, Hongyu Gong, Roman Grundkiewicz, Barry Haddow, Benjamin Hsu, Dávid Javorský, Vera Kloudová, Surafel Lakew, Xutai Ma, Prashant Mathur, Paul McNamee, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Jan Niehues, Xing Niu, John Ortega, Juan Pino, Elizabeth Salesky, Jiatong Shi, Matthias Sperber, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Yogesh Virkar, Alexander Waibel, Changhan Wang, and Shinji Watanabe. 2022. Findings of the IWSLT 2022 evaluation campaign. In Proceedings of the 19th International Conference on Spoken Language Translation (IWSLT 2022), pages 98–157, Dublin, Ireland (in-person and online).
  2. 2.Antonios Anastasopoulos, Ondřej Bojar, Jacob Bremerman, Roldano Cattoni, Maha Elbayad, Marcello Federico, Xutai Ma, Satoshi Nakamura, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Alex Waibel, Changhan Wang, and Matthew Wiesner. 2021. Findings of the IWSLT 2021 Evaluation Campaign. In Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021), Online.
  3. 3.Andrei Andrusenko, Rauf Nasretdinov, and Aleksei Romanenko. 2022. Uconv-conformer: High reduction of input sequence length for end-to-end speech recognition. arXiv preprint arXiv:2208.07657.
  4. 4.Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, Fei Huang, Kevin Knight, Xutai Ma, Ajay Nagesh, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Xing Shi, Sebastian Stüker, Marco Turchi, Alexander Waibel, and Changhan Wang. 2020. FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN. In Proceedings of the 17th International Conference on Spoken Language Translation, pages 1–34, Online.
  5. 5.Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philémon Brakel, and Yoshua Bengio. 2016. End-to-end attention-based large vocabulary speech recognition. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4945–4949.
  6. 6.Maximiliana Behnke and Kenneth Heafield. 2020. Losing heads in the lottery: Pruning transformer attention in neural machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2664–2674, Online.
  7. 7.Maxime Burchi and Valentin Vielzeuf. 2021. Efficient conformer: Progressive downsampling and grouped attention for automatic speech recognition. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 8–15.
  8. 8.Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016. Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation. In NIPS Workshop on end-to-end learning for speech and audio processing, Barcelona, Spain.
  9. 9.Roldano Cattoni, Mattia Antonino Di Gangi, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021. Must-c: A multilingual corpus for end-to-end speech translation. Computer Speech & Language, 66:101155.
  10. 10.William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4960–4964.
  11. 11.Chih-Chiang Chang and Hung-Yi Lee. 2022. Exploring Continuous Integrate-and-Fire for Adaptive Simultaneous Speech Translation. In Proc. Interspeech 2022, pages 5175–5179.
  12. 12.Xuankai Chang, Aswin Shanmugam Subramanian, Pengcheng Guo, Shinji Watanabe, Yuya Fujita, and Motoi Omachi. 2020. End-to-end asr with adaptive span self-attention. In INTERSPEECH.
  13. 13.Junkun Chen, Mingbo Ma, Renjie Zheng, and Liang Huang. 2021. Direct simultaneous speech-to-text translation assisted by synchronized streaming ASR. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4618–4624, Online.
  14. 14.Yun Chen, Yang Liu, Guanhua Chen, Xin Jiang, and Qun Liu. 2020. Accurate word alignment induction from neural machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 566–576, Online.
  15. 15.Colin Cherry and George Foster. 2019. Thinking slow about latency evaluation for simultaneous machine translation.
  16. 16.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy.
  17. 17.Mattia A. Di Gangi, Marco Gaido, Matteo Negri, and Marco Turchi. 2020. On Target Segmentation for Direct Speech Translation. In Proceedings of the 14th Conference of the Association for Machine Translation in the Americas (AMTA 2020), pages 137–150, Virtual.
  18. 18.Javier Ferrando, Gerard I Gállego, Belen Alastruey, Carlos Escolano, and Marta R Costa-jussà. 2022. Towards opening the black box of neural machine translation: Source and target interpretations of the transformer. arXiv e-prints, pages arXiv–2205.
  19. 19.Marco Gaido, Mauro Cettolo, Matteo Negri, and Marco Turchi. 2021a. CTC-based compression for direct speech translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 690–696, Online.
  20. 20.Marco Gaido, Mattia A. Di Gangi, Matteo Negri, and Marco Turchi. 2021b. On Knowledge Distillation for Direct Speech Translation . In Proceedings of CLiC-IT 2020, Online.
  21. 21.Marco Gaido, Matteo Negri, and Marco Turchi. 2022a. Direct speech-to-text translation models as students of text-to-text models. Italian Journal of Computational Linguistics.
  22. 22.Marco Gaido, Sara Papi, Dennis Fucci, Giuseppe Fiameni, Matteo Negri, and Marco Turchi. 2022b. Efficient yet competitive speech translation: FBK@IWSLT2022. In Proceedings of the 19th International Conference on Spoken Language Translation (IWSLT 2022), pages 177–189, Dublin, Ireland (in-person and online).
  23. 23.Sarthak Garg, Stephan Peitz, Udhyakumar Nallasamy, and Matthias Paulik. 2019. Jointly learning to align and translate with transformer models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4453–4462, Hong Kong, China.
  24. 24.Hongyu Gong, Yun Tang, Juan Pino, and Xian Li. 2021. Pay better attention to attention: Head selection in multilingual and multi-domain sequence modeling. Advances in Neural Information Processing Systems, 34:2668–2681.
  25. 25.Alex Graves. 2012. Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711.
  26. 26.Alex Graves. 2013. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850.
  27. 27.Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. In Proceedings of the 23rd international conference on Machine learning (ICML), pages 369–376, Pittsburgh, Pennsylvania.
  28. 28.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. Conformer: Convolution-augmented Transformer for Speech Recognition. In Proc. Interspeech 2020, pages 5036–5040.
  29. 29.Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang. 2021. Recent developments on espnet toolkit boosted by conformer. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5874–5878.
  30. 30.Hou Jeung Han, Mohd Abbas Zaidi, Sathish Reddy Indurthi, Nikhil Kumar Lakumarapu, Beomseok Lee, and Sangha Kim. 2020. End-to-end simultaneous translation system for IWSLT2020 using modality agnostic meta-learning. In Proceedings of the 17th International Conference on Spoken Language Translation, pages 62–68, Online.
  31. 31.Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019. Do attention heads in bert track syntactic dependencies? arXiv preprint arXiv:1911.12246.
  32. 32.Jae-young Jo and Sung-Hyon Myaeng. 2020. Roles and utilization of attention heads in transformer-based neural language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3404–3417, Online.
  33. 33.Alina Karakanta, Sara Papi, Matteo Negri, and Marco Turchi. 2021. Simultaneous speech translation for live subtitling: from delay to display. In Proceedings of the 1st Workshop on Automatic Spoken Language Translation in Real-World Settings (ASLTRW), pages 35–48, Virtual.
  34. 34.Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer. 2022. Squeezeformer: An efficient transformer for automatic speech recognition. arxiv:2206.00888.
  35. 35.Yoon Kim and Alexander M. Rush. 2016. Sequence-Level Knowledge Distillation. In Proc. of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1317–1327, Austin, Texas.
  36. 36.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  37. 37.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7057–7075, Online.
  38. 38.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4365–4374, Hong Kong, China.
  39. 39.Mathis Lamarre, Catherine Chen, and Fatma Deniz. 2022. Attention weights accurately predict language representations in the brain. bioRxiv.
  40. 40.Dan Liu, Mengge Du, Xiaoxi Li, Yuchen Hu, and Lirong Dai. 2021a. The USTC-NELSLIP systems for simultaneous speech translation task at IWSLT 2021. In Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021), pages 30–38, Bangkok, Thailand (online).
  41. 41.Dan Liu, Mengge Du, Xiaoxi Li, Ya Li, and Enhong Chen. 2021b. Cross attention augmented transducer networks for simultaneous translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 39–55, Online and Punta Cana, Dominican Republic.
  42. 42.Danni Liu, Gerasimos Spanakis, and Jan Niehues. 2020. Low-Latency Sequence-to-Sequence Speech Recognition and Translation by Partial Hypothesis Selection. In Proc. Interspeech 2020, pages 3620–3624.
  43. 43.Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019. STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3025–3036, Florence, Italy.
  44. 44.Xutai Ma, Mohammad Javad Dousti, Changhan Wang, Jiatao Gu, and Juan Pino. 2020a. SIMULEVAL: An evaluation toolkit for simultaneous translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 144–150, Online.
  45. 45.Xutai Ma, Juan Pino, and Philipp Koehn. 2020b. SimulMT to SimulST: Adapting simultaneous text translation to end-to-end simultaneous speech translation. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 582–587, Suzhou, China.
  46. 46.Xutai Ma, Yongqiang Wang, Mohammad Javad Dousti, Philipp Koehn, and Juan Pino. 2021. Streaming simultaneous speech translation with augmented memory transformer. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7523–7527. IEEE.
  47. 47.Ha Nguyen, Yannick Estève, and Laurent Besacier. 2021. An empirical study of end-to-end simultaneous speech translation decoding strategies. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7528–7532. IEEE.
  48. 48.Motoi Omachi, Brian Yan, Siddharth Dalmia, Yuya Fujita, and Shinji Watanabe. 2022. Align, write, re-order: Explainable end-to-end speech translation via operation sequence generation. arXiv preprint arXiv:2211.05967.
  49. 49.Sara Papi, Marco Gaido, Matteo Negri, and Andrea Pilzer. 2023. Reproducibility is nothing without correctness: The importance of testing code in nlp. ArXiv, abs/2303.16166.
  50. 50.Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2021. Speechformer: Reducing information loss in direct speech translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1698–1706, Online and Punta Cana, Dominican Republic.
  51. 51.Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2022a. Does simultaneous speech translation need simultaneous models? In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 141–153, Abu Dhabi, United Arab Emirates.
  52. 52.Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2022b. Over-generation cannot be rewarded: Length-adaptive average lagging for simultaneous speech translation. In Proceedings of the Third Workshop on Automatic Simultaneous Translation, pages 12–17, Online.
  53. 53.Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019. SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. In Proc. Interspeech 2019, pages 2613–2617.
  54. 54.Peter Polák, Ngoc-Quan Pham, Tuan Nam Nguyen, Danni Liu, Carlos Mullov, Jan Niehues, Ondřej Bojar, and Alexander Waibel. 2022. CUNI-KIT system for simultaneous speech translation task at IWSLT 2022. In Proceedings of the 19th International Conference on Spoken Language Translation (IWSLT 2022), pages 277–285, Dublin, Ireland (in-person and online).
  55. 55.Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels.
  56. 56.Alessandro Raganato and Jörg Tiedemann. 2018. An analysis of encoder representations in transformer-based machine translation. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 287–297, Brussels, Belgium.
  57. 57.Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2020. SimulSpeech: End-to-end simultaneous speech to text translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3787–3796, Online.
  58. 58.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany.
  59. 59.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the Inception Architecture for Computer Vision. In Proc. of 2016 IEEE CVPR, pages 2818–2826, Las Vegas, Nevada, United States.
  60. 60.Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018. An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 26–35, Brussels, Belgium.
  61. 61.Jörg Tiedemann. 2016. OPUS – parallel corpora for everyone. In Proceedings of the 19th Annual Conference of the European Association for Machine Translation: Projects/Products, Riga, Latvia.
  62. 62.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30.
  63. 63.Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63–76, Florence, Italy.
  64. 64.Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020. fairseq s2t: Fast speech-to-text modeling with fairseq. In Proceedings of the 2020 Conference of the Asian Chapter of the Association for Computational Linguistics (AACL): System Demonstrations.
  65. 65.Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017. Sequence-to-Sequence Models Can Directly Translate Foreign Speech. In Proceedings of Interspeech 2017, pages 2625–2629, Stockholm, Sweden.
  66. 66.Mohd Abbas Zaidi, Beomseok Lee, Sangha Kim, and Chanwoo Kim. 2022. Cross-Modal Decision Regularization for Simultaneous Speech Translation. In Proc. Interspeech 2022, pages 116–120.
  67. 67.Mohd Abbas Zaidi, Beomseok Lee, Nikhil Kumar Lakumarapu, Sangha Kim, and Chanwoo Kim. 2021. Decision attentive regularization to improve simultaneous speech translation systems. arXiv preprint arXiv:2110.15729.
  68. 68.Xingshan Zeng, Liangyou Li, and Qun Liu. 2021. RealTranS: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2461–2474, Online.
  69. 69.Thomas Zenkel, Joern Wuebker, and John DeNero. 2019. Adding interpretable attention to neural translation models improves word alignment. arXiv preprint arXiv:1901.11359.
  70. 70.Shaolei Zhang and Yang Feng. 2022. Information-transport-based policy for simultaneous translation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 992–1013, Abu Dhabi, United Arab Emirates.
  71. 71.Baigong Zheng, Kaibo Liu, Renjie Zheng, Mingbo Ma, Hairong Liu, and Liang Huang. 2020. Simultaneous translation policies: From fixed to adaptive. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2847–2853, Online.

Citation

MLA
Papi, S., et al. “Attention as a Guide for Simultaneous Speech Translation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 13340–56, https://doi.org/10.18653/v1/2023.acl-long.745.
APA
Papi, S., Negri, M., & Turchi, M. (2023). Attention as a Guide for Simultaneous Speech Translation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13340–13356. https://doi.org/10.18653/v1/2023.acl-long.745
Chicago
Papi, S., M. Negri, and M. Turchi. 2023. “Attention as a Guide for Simultaneous Speech Translation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13340–56. https://doi.org/10.18653/v1/2023.acl-long.745.
Harvard
Papi, S., Negri, M. and Turchi, M. (2023) “Attention as a Guide for Simultaneous Speech Translation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13340–13356. Available at: https://doi.org/10.18653/v1/2023.acl-long.745.
Vancouver
1. Papi S, Negri M, Turchi M (2023) Attention as a Guide for Simultaneous Speech Translation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13340–13356

BibTeX

@inproceedings{papi-etal-2023-attention,
    title = "Attention as a Guide for Simultaneous Speech Translation",
    author = "Papi, Sara  and
      Negri, Matteo  and
      Turchi, Marco",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.745/",
    doi = "10.18653/v1/2023.acl-long.745",
    pages = "13340--13356"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/