Quality-Aware Decoding for Neural Machine Translation

Patrick FernandesAntónio FarinhasRicardo ReiJosé Guilherme Camargo de SouzaPerez OgayoGraham NeubigAndré F. T. Martins

article2022NAACL65 citations

Proposes a unified quality-aware decoding framework that integrates modern reference-free and reference-based metrics into candidate reranking and Minimum Bayes Risk decoding, substantially outperforming standard beam search across modern neural translation benchmarks and human assessments.

Listen

Standard neural machine translation systems typically generate translations by searching for the single most probable sequence of words using beam search. While this probability-based approach is standard, high sequence likelihood frequently fails to correlate with actual translation quality, resulting in awkward or inaccurate text. In recent years, automated quality evaluation metrics have improved dramatically, yet these modern evaluators are rarely integrated into the translation generation process itself. The article systematically evaluates whether and how modern quality evaluation metrics can be embedded directly into decoding to produce superior translations.

The researchers propose a unified framework termed quality-aware decoding, decoupling the generation process into candidate generation followed by candidate ranking. The investigation comprehensively evaluates two main ranking mechanisms across four translation benchmarks and two model scales: candidate reranking using quality estimation metrics, and minimum Bayes risk decoding, which selects the candidate maximizing expected quality against all others. Additionally, the article tests a two-stage hybrid approach that pre-filters candidates via reranking before applying minimum Bayes risk decoding. Candidate generation strategies—including standard beam search, vanilla random sampling, and nucleus sampling—are also contrasted alongside training variations like label smoothing.

The empirical findings demonstrate several critical insights for translation system design. First, quality-aware decoding systematically outperforms standard beam search across modern automated metrics and rigorous human evaluations. Second, tuned multi-metric reranking and the two-stage hybrid method consistently achieve the strongest overall performance; for English-to-Russian translation, these approaches cut critical and major lexical selection errors by roughly half compared to the baseline. Third, candidate generation strategy heavily influences downstream success: nucleus sampling and beam search scale well with larger candidate pools, whereas vanilla sampling performs poorly unless specific training regularizations like label smoothing are removed. Finally, the study reveals a notable risk of metric overfitting: optimizing exclusively for a single quality estimation metric produces inflated automated scores while degrading actual human-judged quality, whereas tuned multi-metric approaches remain robust.

These findings indicate that organizations relying on automated translation can substantially improve translation accuracy, grammatical register, and fluency without retraining underlying base models. However, advanced ranking introduces significant computational overhead, especially with minimum Bayes risk decoding. Organizations seeking immediate deployment should prioritize tuned multi-metric reranking or the two-stage hybrid pipeline using candidate pools generated via nucleus sampling or beam search, as these deliver the best trade-off between computational cost and translation quality. Further engineering efforts should focus on metric efficiency, caching mechanisms, and model distillation to lower inference latency for production environments.

Fernandes et al (2022).pdf
Cover for Quality-Aware Decoding for Neural Machine Translation

Table of Contents

  • 1 Introduction
  • 2 Candidate Generation and Ranking
  • 2.1 Candidate Generation
  • 2.2 Ranking
  • 2.2.1 N-best Reranking
  • 2.2.2 Minimum Bayes Risk (MBR) Decoding
  • 3 Quality-Aware Decoding
  • 3.1 Reference-based Metrics
  • 3.2 Reference-free Metrics
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Results
  • 4.2.1 Impact of Candidate Generation
  • 4.2.2 Impact of Label Smoothing
  • 4.2.3 Impact of Ranking and Metrics
  • 4.2.4 Human Evaluation
  • 4.2.5 Improved Human Evaluation
  • 5 Related Work
  • 6 Conclusions and Future Work
  • Acknowledgments
  • References
  • Supplemental Material
  • A Training Details
  • B Additional Results
  • C Human Study
  • D MQM Framework

Knowls

  1. Knowl 1 — Quality-aware decoding separates candidate generation from quality-based selection

    model/method

    Quality-aware decoding for neural machine translation (NMT) treats decoding as two stages. An NMT model with parameters θ\theta first generates a finite set C\mathcal{C} of candidate translations for source sentence xx; a separate ranking procedure then selects the final translation from C\mathcal{C} using one or more translation-quality metrics. This contrasts with maximum-a-posteriori decoding, which selects the candidate with the greatest model probability pθ(y∣x)p_\theta(y\mid x). The ranking stage can be paired with beam search or sampling and does not require changing the NMT model's training procedure.

  2. Knowl 2 — Fixed and tuned N-best reranking use reference-free metric features

    model/method

    For a source sentence xx and a candidate set C\mathcal{C} of NN translations, fixed N-best reranking selects the candidate with the highest score from one reference-free quality-estimation (QE) metric ff. Tuned N-best reranking instead uses KK candidate-scoring features fkf_k, including the QE metrics and the NMT model log-likelihood, and selects by their weighted sum:

    y^fixed=arg⁡max⁡y∈Cf(y),y^tuned=arg⁡max⁡y∈C∑k=1Kwkfk(y).\hat y_{\mathrm{fixed}}=\arg\max_{y\in\mathcal{C}} f(y), \qquad \hat y_{\mathrm{tuned}}=\arg\max_{y\in\mathcal{C}}\sum_{k=1}^{K}w_k f_k(y).

    Here, yy is a candidate translation, wkw_k is the weight for feature fkf_k, and the weights are tuned with minimum error rate training (MERT) on validation data to maximize a chosen reference-based metric. That reference-based metric is used for tuning only; at test time, tuned reranking uses the learned weights and reference-free features. If evaluating QE feature ii costs CiC_i, fixed reranking costs O(NCf)O(NC_f) for its one metric, while tuned reranking costs O(N∑iCi)O(N\sum_i C_i) over its QE features.

  3. Knowl 3 — Minimum Bayes risk decoding selects candidates by expected reference-based utility

    model/method

    Minimum Bayes risk (MBR) decoding chooses the candidate with the greatest expected utility under the model's translation distribution, rather than the candidate with the greatest model probability. For source sentence xx, candidate set C\mathcal{C}, and reference-based utility metric uu, the expected utility of candidate yy is estimated from MM translations sampled from the model:

    y^MBR=arg⁡max⁡y∈C1M∑j=1Mu(z(j),y),z(j)∼pθ(⋅∣x).\hat y_{\mathrm{MBR}}=\arg\max_{y\in\mathcal{C}}\frac{1}{M}\sum_{j=1}^{M}u(z^{(j)},y), \qquad z^{(j)}\sim p_\theta(\cdot\mid x).

    Here, z(j)z^{(j)} is a sampled translation treated as a reference for comparing candidate yy, and pθ(⋅∣x)p_\theta(\cdot\mid x) is the NMT model's distribution over translations given xx. With NN candidates and an equally sized set of model samples, all pairwise comparisons cost O(N2Cu)O(N^2 C_u), where CuC_u is the cost of evaluating the utility metric once. Replacing model samples with candidates from beam search or nucleus sampling yields a biased approximation to the expectation.

  4. Knowl 4 — A tuned-reranking shortlist makes MBR decoding less expensive

    model/method

    The two-stage tuned-reranking-to-MBR method first scores all NN generated candidates with a tuned N-best reranker, retains the top LL candidates, and applies MBR decoding to that shortlist using a reference-based metric. The final translation is the shortlist candidate with the greatest average pairwise utility against the shortlist's translations. If the first-stage QE metrics have evaluation costs CiC_i and the MBR metric has cost CuC_u, the total cost is O(N∑iCi+L2Cu)O(N\sum_i C_i+L^2C_u), compared with quadratic MBR cost over the full candidate set. In the experiments, this pruning made it possible to use MBR on more promising candidates while controlling its computational cost.

  5. Knowl 5 — The ranking framework combines learned reference-based and reference-free metrics

    model/method

    The study uses BLEURT-20 and the WMT20 COMET model as learned reference-based metrics. BLEURT-20 is based on RemBERT and trained to predict human direct-assessment scores; the COMET model is based on XLM-R and uses the source sentence and a reference to score a translation. These metrics serve as MBR utilities, tuning objectives, and automatic evaluation measures. The reference-free QE features are COMET-QE, TransQuest, MBART-QE, and OpenKiwi-MQM: the first three predict direct-assessment-related quality scores, while OpenKiwi-MQM predicts multidimensional quality metric (MQM) annotations. BLEU and chrF are also used for automatic evaluation and, in some experiments, as reranking or MBR objectives.

  6. Knowl 6 — Experiments compare two model scales across four translation directions

    experimental setup

    The experiments cover four English-to-target-language directions in two settings. The large setting uses non-ensembled WMT19 news-translation Transformer models for English–German and English–Russian, with 6 layers, 16 attention heads, 1024-dimensional embeddings, and 8192-dimensional hidden layers; newstest19 is used for validation and newstest20 for testing. The small setting trains 6-layer Transformers from scratch on IWSLT17 English–German and English–French data, each with slightly more than 200,000 training examples; the models have 4 attention heads, 512-dimensional embeddings, and 1024-dimensional hidden layers. Small-model training uses a joint 20,000-unit SentencePiece vocabulary, Adam with β1=0.9\beta_1=0.9 and β2=0.98\beta_2=0.98, an inverse-square-root learning-rate schedule with initial rate 5×10−45\times10^{-4}, and 4,000 warm-up steps; label smoothing is 0.1 when enabled. Beam search with beam size 5 is the MAP-based baseline. N-best reranking uses up to 200 candidates, while the main MBR comparisons use 50 candidates; these settings were reported to have roughly comparable runtime for tuned reranking and MBR.

  7. Knowl 7 — Candidate-generation performance depends on search method and label smoothing

    empirical result

    The candidate-generation comparison tests beam search, vanilla sampling, and nucleus sampling, using nucleus threshold p=0.6p=0.6. Increasing the candidate count generally improves reranking and MBR results, especially when candidates are generated by vanilla sampling, but vanilla-sampling rankers often remain below the beam-size-5 baseline. Beam- and nucleus-generated candidates are more consistently competitive: they generally match or exceed the baseline on BLEU and produce substantial COMET gains. With the large models, lexical-metric performance can fall as the candidate set grows; nucleus sampling has an advantage over beam search for COMET in that setting. The later experiments therefore use nucleus sampling for large models and beam search for small models. In a small-model comparison, removing label smoothing improves vanilla-sampling results but worsens results with nucleus sampling; even without label smoothing, vanilla sampling remains uncompetitive with beam search or nucleus sampling.

  8. Knowl 8 — Automatic evaluations favor quality-aware ranking on learned metrics

    data/table

    Across the English–German comparisons, quality-aware methods improve learned-metric scores, particularly COMET, although gains on learned metrics do not always accompany gains on lexical metrics. Selected reported values are BLEU and COMET:

    • Large WMT20 English–German: beam-size-5 baseline, 36.01 and 0.5795; tuned reranking optimized for COMET, 34.26 and 0.6276; MBR with COMET, 33.04 and 0.6359; tuned reranking followed by MBR with COMET, 34.20 and 0.6418.
    • Small IWSLT17 English–German: baseline, 29.12 and 0.3028; tuned reranking optimized for COMET, 30.16 and 0.4721; MBR with COMET, 29.43 and 0.4480; two-stage tuned reranking plus MBR with COMET, 29.46 and 0.5005.

    For the large model, the two-stage COMET system reaches the highest COMET value among these English–German systems, but not the highest BLEU. For the small model, the two-stage system also has the highest COMET among these methods, while tuned reranking optimized for BLEU reaches BLEU 30.51. The broader comparisons across the four language directions show the same central qualification: optimizing or selecting by a learned metric can improve that metric without improving lexical metrics.

  9. Knowl 9 — Scalar human judgments expose metric overfitting but favor several quality-aware systems

    empirical result

    Professional native speakers of each target language rated translations on a 1–5 scale, where 1 indicates no overlap in meaning and 5 indicates equivalent meaning expressed naturally. The evaluation sampled 300 sentences per language direction, deduplicated each source sentence's hypotheses, and presented them in randomized side-by-side order. Mean human ratings for baseline, fixed reranking with COMET-QE, tuned reranking with COMET, MBR with COMET, and the two-stage COMET system were, respectively:

    • WMT20 English–German: 4.28, 4.19, 4.33, 4.27, 4.30.
    • WMT20 English–Russian: 3.62, 3.25, 3.65, 3.66, 3.72; the two-stage system was significantly better than baseline at p<0.05p<0.05.
    • IWSLT17 English–German: 3.68, 3.67, 3.90, 3.79, 3.83; tuned reranking, MBR, and the two-stage system were significantly better than baseline at p<0.05p<0.05.
    • IWSLT17 English–French: 3.92, 3.63, 4.05, 4.05, 4.09; all three methods other than fixed reranking were significantly better than baseline at p<0.05p<0.05.

    Fixed reranking with COMET-QE achieved higher COMET than baseline in all four directions but received lower human ratings in three of them, illustrating that optimizing a learned metric can make it a less reliable system-level indicator. The tuned, MBR, and combined methods generally received ratings at least as high as baseline, with the combined system highest in two of the four directions.

  10. Knowl 10 — MQM annotations show the strongest gains for tuned and two-stage decoding

    data/table

    For the large WMT20 models, expert annotators marked translation errors by severity; the reported MQM score uses severity weights of 1 for minor, 5 for major, and 10 for critical errors. Each entry below gives minor, major, and critical error counts followed by the MQM score:

    • English–German: reference 24, 67, 0; 97.04. Baseline 8, 139, 0; 95.66. Fixed COMET-QE reranking 15, 204, 0; 93.47. Tuned COMET reranking 12, 109, 0; 96.20. MBR with COMET 11, 161, 0; 94.38. Two-stage tuned reranking plus MBR with COMET 10, 138, 0; 95.44.
    • English–Russian: reference 5, 11, 0; 99.30. Baseline 17, 239, 49; 79.78. Fixed COMET-QE reranking 13, 254, 80; 76.25. Tuned COMET reranking 9, 141, 45; 85.97. MBR with COMET 8, 182, 40; 83.65. Two-stage tuned reranking plus MBR with COMET 11, 134, 45; 86.78.

    Tuned reranking has the highest English–German MQM score, while the two-stage method has the highest English–Russian score; both are significantly better than the corresponding baseline at p<0.05p<0.05. Fixed COMET-QE reranking has lower MQM scores than baseline in both directions despite its strong automatic COMET scores. Error-category analysis further reports that tuned and two-stage methods reduce grammatical-register errors in English–German while increasing lexical-selection errors; in English–Russian, their lexical-selection errors are approximately half the baseline count.

Coverage note — No substantial contributed result was deliberately omitted; detailed training and annotation procedures are included where they support interpretation, while individual appendix error-category plots are summarized by their reported patterns.

References

  1. 1.Chantal Amrhein and Rico Sennrich. 2022. Identifying weaknesses in machine translation metrics through minimum bayes risk decoding: A case study for comet.
  2. 2.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 1–61, Florence, Italy. Association for Computational Linguistics.
  3. 3.Sumanta Bhattacharyya, Amirmohammad Rooshenas, Subhajit Naskar, Simeng Sun, Mohit Iyyer, and Andrew McCallum. 2021. Energy-based reranking: Improving neural machine translation using energy-based models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4528–4537, Online. Association for Computational Linguistics.
  4. 4.Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012. WIT3: Web inventory of transcribed and translated talks. In Proceedings of the 16th Annual conference of the European Association for Machine Translation, pages 261–268, Trento, Italy. European Association for Machine Translation.
  5. 5.Hyung Won Chung, Thibault Fevry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2021. Rethinking embedding coupling in pre-trained language models. In Tenth International Conference on Learning Representations, ICLR.
  6. 6.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  7. 7.Kevin Duh and Katrin Kirchhoff. 2008. Beyond log-linear models: Boosted minimum error rate training for n-best re-ranking. In Proceedings of ACL-08: HLT, Short Papers, pages 37–40, Columbus, Ohio. Association for Computational Linguistics.
  8. 8.Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018. Classical structured prediction losses for sequence to sequence learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 355–364, New Orleans, Louisiana. Association for Computational Linguistics.
  9. 9.Bryan Eikema and Wilker Aziz. 2020. Is MAP decoding all you need? the inadequacy of the mode in neural machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4506–4520, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  10. 10.Bryan Eikema and Wilker Aziz. 2021. Sampling-based minimum bayes risk decoding for neural machine translation.
  11. 11.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  12. 12.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  13. 13.Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2021b. Minimum bayes risk decoding with neural metrics of translation quality.
  14. 14.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondrej Bojar. 2021c. Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain. In Proceedings of the Sixth Conference on Machine Translation (WMT), pages 716–757. NRC.
  15. 15.Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013. Continuous measurement scales in human evaluation of machine translation. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 33–41, Sofia, Bulgaria. Association for Computational Linguistics.
  16. 16.Alex Graves. 2012. Sequence transduction with recurrent neural networks.
  17. 17.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In Eighth International Conference on Learning Representations, ICLR.
  18. 18.Fabio Kepler, Jonay Trénous, Marcos Treviso, Miguel Vera, and André F. T. Martins. 2019. OpenKiwi: An open source framework for quality estimation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 117–122, Florence, Italy. Association for Computational Linguistics.
  19. 19.Samuel Kiegeland and Julia Kreutzer. 2021. Revisiting the weaknesses of reinforcement learning for neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1673–1681, Online. Association for Computational Linguistics.
  20. 20.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Third International Conference on Learning Representations, ICLR.
  21. 21.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation. In Proceedings of the Sixth Conference on Machine Translation, pages 478–494, Online. Association for Computational Linguistics.
  22. 22.Philipp Koehn and Rebecca Knowles. 2017. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 28–39, Vancouver. Association for Computational Linguistics.
  23. 23.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  24. 24.Shankar Kumar and William Byrne. 2002. Minimum bayes-risk word alignments of bilingual texts. In Proceedings of the ACL-02 Conference on Empirical Methods in Natural Language Processing - Volume 10, EMNLP ’02, page 140–147, USA. Association for Computational Linguistics.
  25. 25.Shankar Kumar and William Byrne. 2004. Minimum Bayes-risk decoding for statistical machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 169–176, Boston, Massachusetts, USA. Association for Computational Linguistics.
  26. 26.Alon Lavie and Michael J. Denkowski. 2009. The meteor metric for automatic evaluation of machine translation. Machine Translation, 23(2-3):105–115.
  27. 27.Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Lespiau Jean-Baptiste, Ioannis Antonoglou, Karen Simonyan, and Oriol Vinyals. 2021. Machine translation decoding beyond beam search. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8410–8434, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  28. 28.Ann Lee, Michael Auli, and Marc’Aurelio Ranzato. 2021. Discriminative reranking for neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7250–7264, Online. Association for Computational Linguistics.
  29. 29.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  30. 30.Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014. Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics. Tradumàtica: tecnologies de la traducció, 0:455–463.
  31. 31.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020a. Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4984–4997, Online. Association for Computational Linguistics.
  32. 32.Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020b. Results of the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 688–725, Online. Association for Computational Linguistics.
  33. 33.Clara Meister, Ryan Cotterell, and Tim Vieira. 2020. If beam search is the answer, what was the question? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2173–2185, Online. Association for Computational Linguistics.
  34. 34.Mathias Müller and Rico Sennrich. 2021. Understanding the properties of minimum Bayes risk decoding in neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 259–272, Online. Association for Computational Linguistics.
  35. 35.Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. 2019. When does label smoothing help? In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  36. 36.Kenton Murray and David Chiang. 2018. Correcting length bias in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 212–223, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Graham Neubig. 2013. Travatar: A forest-to-string machine translation engine based on tree transducers. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 91–96, Sofia, Bulgaria. Association for Computational Linguistics.
  38. 38.Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019. Facebook FAIR’s WMT19 news translation task submission. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 314–319, Florence, Italy. Association for Computational Linguistics.
  39. 39.Franz Josef Och. 2003. Minimum error rate training in statistical machine translation. In Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, pages 160–167, Sapporo, Japan. Association for Computational Linguistics.
  40. 40.Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato. 2018. Analyzing uncertainty in neural machine translation. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 3956–3965. PMLR.
  41. 41.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  42. 42.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  43. 43.Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  44. 44.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels. Association for Computational Linguistics.
  45. 45.Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021. Learning compact metrics for MT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 751–762, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov. 2020. TransQuest: Translation quality estimation with cross-lingual transformers. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5070–5081, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  47. 47.Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André F. T. Martins, and Alon Lavie. 2021. Are references really needed? unbabel-IST 2021 submission for the metrics shared task. In Proceedings of the Sixth Conference on Machine Translation, pages 1030–1040, Online. Association for Computational Linguistics.
  48. 48.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020a. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  49. 49.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020b. Unbabel’s participation in the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 911–920, Online. Association for Computational Linguistics.
  50. 50.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  51. 51.Uri Shaham and Omer Levy. 2021. What do you get when you cross beam search with nucleus sampling?
  52. 52.Libin Shen, Anoop Sarkar, and Franz Josef Och. 2004. Discriminative reranking for machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 177–184, Boston, Massachusetts, USA. Association for Computational Linguistics.
  53. 53.Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016. Minimum risk training for neural machine translation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1683–1692, Berlin, Germany. Association for Computational Linguistics.
  54. 54.Raphael Shu and Hideki Nakayama. 2017. Later-stage minimum bayes-risk decoding for neural machine translation.
  55. 55.David A. Smith and Jason Eisner. 2006. Minimum risk annealing for training log-linear models. In Proceedings of the COLING/ACL 2006 Main Conference Poster Sessions, pages 787–794, Sydney, Australia. Association for Computational Linguistics.
  56. 56.Lucia Specia, Frédéric Blain, Marina Fomicheva, Erick Fonseca, Vishrav Chaudhary, Francisco Guzmán, and André F. T. Martins. 2020. Findings of the WMT 2020 shared task on quality estimation. In Proceedings of the Fifth Conference on Machine Translation, pages 743–764, Online. Association for Computational Linguistics.
  57. 57.Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021. Findings of the WMT 2021 shared task on quality estimation. In Proceedings of the Sixth Conference on Machine Translation, pages 684–725, Online. Association for Computational Linguistics.
  58. 58.Felix Stahlberg and Bill Byrne. 2019. On NMT search errors and model errors: Cat got your tongue? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3356–3362, Hong Kong, China. Association for Computational Linguistics.
  59. 59.Felix Stahlberg, Adrià de Gispert, Eva Hasler, and Bill Byrne. 2017. Neural machine translation by minimising the Bayes-risk with respect to syntactic translation lattices. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 362–368, Valencia, Spain. Association for Computational Linguistics.
  60. 60.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc.
  61. 61.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826.
  62. 62.Chaojun Wang and Rico Sennrich. 2020. On exposure bias, hallucination and domain shift in neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3544–3552, Online. Association for Computational Linguistics.
  63. 63.John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019. Beyond BLEU:training neural machine translation with semantic similarity. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4344–4355, Florence, Italy. Association for Computational Linguistics.
  64. 64.Yilin Yang, Liang Huang, and Mingbo Ma. 2018. Breaking the beam search curse: A study of (re-)scoring methods and stopping criteria for neural machine translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3054–3059, Brussels, Belgium. Association for Computational Linguistics.
  65. 65.Kyra Yee, Yann Dauphin, and Michael Auli. 2019. Simple and effective noisy channel modeling for neural machine translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5696–5701, Hong Kong, China. Association for Computational Linguistics.
  66. 66.Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomás Kociský. 2017. The neural noisy channel. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  67. 67.Chrysoula Zerva, Daan van Stigt, Ricardo Rei, Ana C Farinha, Pedro Ramos, José G. C. de Souza, Taisiya Glushkova, Miguel Vera, Fabio Kepler, and André F. T. Martins. 2021. IST-unbabel 2021 submission for the quality estimation shared task. In Proceedings of the Sixth Conference on Machine Translation, pages 961–972, Online. Association for Computational Linguistics.
  68. 68.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.

Citation

MLA
Fernandes, P., et al. “Quality-Aware Decoding for Neural Machine Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 1396–412, https://doi.org/10.18653/v1/2022.naacl-main.100.
APA
Fernandes, P., Farinhas, A., Rei, R., Souza, J. G. C. de ., Ogayo, P., Neubig, G., & Martins, A. F. T. (2022). Quality-Aware Decoding for Neural Machine Translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1396–1412. https://doi.org/10.18653/v1/2022.naacl-main.100
Chicago
Fernandes, P., A. Farinhas, R. Rei, et al. 2022. “Quality-Aware Decoding for Neural Machine Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1396–1412. https://doi.org/10.18653/v1/2022.naacl-main.100.
Harvard
Fernandes, P. et al. (2022) “Quality-Aware Decoding for Neural Machine Translation”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 1396–1412. Available at: https://doi.org/10.18653/v1/2022.naacl-main.100.
Vancouver
1. Fernandes P, Farinhas A, Rei R, Souza JGC de, Ogayo P, Neubig G, Martins AFT (2022) Quality-Aware Decoding for Neural Machine Translation. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 1396–1412

BibTeX

@inproceedings{fernandes-etal-2022-quality,
    title = "Quality-Aware Decoding for Neural Machine Translation",
    author = "Fernandes, Patrick  and
      Farinhas, Ant{\'o}nio  and
      Rei, Ricardo  and
      C. de Souza, Jos{\'e} G.  and
      Ogayo, Perez  and
      Neubig, Graham  and
      Martins, Andre",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.100/",
    doi = "10.18653/v1/2022.naacl-main.100",
    pages = "1396--1412"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/