Quality-Aware Decoding for Neural Machine Translation
Patrick FernandesAntónio FarinhasRicardo ReiJosé Guilherme Camargo de SouzaPerez OgayoGraham NeubigAndré F. T. Martins
Proposes a unified quality-aware decoding framework that integrates modern reference-free and reference-based metrics into candidate reranking and Minimum Bayes Risk decoding, substantially outperforming standard beam search across modern neural translation benchmarks and human assessments.
Standard neural machine translation systems typically generate translations by searching for the single most probable sequence of words using beam search. While this probability-based approach is standard, high sequence likelihood frequently fails to correlate with actual translation quality, resulting in awkward or inaccurate text. In recent years, automated quality evaluation metrics have improved dramatically, yet these modern evaluators are rarely integrated into the translation generation process itself. The article systematically evaluates whether and how modern quality evaluation metrics can be embedded directly into decoding to produce superior translations.
The researchers propose a unified framework termed quality-aware decoding, decoupling the generation process into candidate generation followed by candidate ranking. The investigation comprehensively evaluates two main ranking mechanisms across four translation benchmarks and two model scales: candidate reranking using quality estimation metrics, and minimum Bayes risk decoding, which selects the candidate maximizing expected quality against all others. Additionally, the article tests a two-stage hybrid approach that pre-filters candidates via reranking before applying minimum Bayes risk decoding. Candidate generation strategies—including standard beam search, vanilla random sampling, and nucleus sampling—are also contrasted alongside training variations like label smoothing.
The empirical findings demonstrate several critical insights for translation system design. First, quality-aware decoding systematically outperforms standard beam search across modern automated metrics and rigorous human evaluations. Second, tuned multi-metric reranking and the two-stage hybrid method consistently achieve the strongest overall performance; for English-to-Russian translation, these approaches cut critical and major lexical selection errors by roughly half compared to the baseline. Third, candidate generation strategy heavily influences downstream success: nucleus sampling and beam search scale well with larger candidate pools, whereas vanilla sampling performs poorly unless specific training regularizations like label smoothing are removed. Finally, the study reveals a notable risk of metric overfitting: optimizing exclusively for a single quality estimation metric produces inflated automated scores while degrading actual human-judged quality, whereas tuned multi-metric approaches remain robust.
These findings indicate that organizations relying on automated translation can substantially improve translation accuracy, grammatical register, and fluency without retraining underlying base models. However, advanced ranking introduces significant computational overhead, especially with minimum Bayes risk decoding. Organizations seeking immediate deployment should prioritize tuned multi-metric reranking or the two-stage hybrid pipeline using candidate pools generated via nucleus sampling or beam search, as these deliver the best trade-off between computational cost and translation quality. Further engineering efforts should focus on metric efficiency, caching mechanisms, and model distillation to lower inference latency for production environments.
- Paper: COMET: A Neural Framework for MT Evaluation, Ricardo Rei et al. (2020). COMET introduces the neural translation-quality estimators that make the source’s metric-based candidate ranking possible.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). Nucleus sampling explains a candidate-generation strategy the source tests and finds effective when paired with quality-aware ranking.
- Paper: Bleu: a Method for Automatic Evaluation of Machine Translation, Kishore Papineni et al. (2002). BLEU establishes the automatic translation-evaluation tradition whose limits motivate ranking candidates by richer quality metrics.
- Paper: On the Limitations of Reference-Free Evaluations of Generated Text, Daniel Deutsch et al. (2022). By optimizing reference-free metrics during decoding, this study stress-tests the source’s quality-aware strategy and shows how metric gains can conceal poor actual output quality.
