Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies

Tom KocmiVilém ZouharChristian FedermannMatt Post

article2024ACL53 citations

Establishes practical score difference thresholds across modern machine translation metrics to determine when system-level score gains reflect perceivable translation quality improvements to human evaluators.

Listen

Modern machine translation research and deployment rely heavily on automated scoring metrics to evaluate model improvements. Historically, the field relied on a single standard metric, BLEU, but modern neural metrics have replaced it with a fragmented landscape of competing measurement tools. These newer metrics operate on vastly different numerical score scales, making it difficult for decision-makers and practitioners to understand what score difference constitutes a real, meaningful quality improvement that human users would notice.

The article aims to resolve this ambiguity by determining the score differences required for various metrics to achieve system-level improvements that align with human judgment. It establishes a unified framework that maps raw metric score differences directly to estimated human pairwise ranking accuracy.

To conduct this evaluation, the authors analyzed ToShip23, an extensive human evaluation dataset containing over 3 million segments and 6,530 system pairs across 94 languages and more than 10 domains. By grouping system score deltas into bins and fitting mathematical curves to the empirical data, the study established reliable thresholds of human agreement. Key findings were subsequently validated on public benchmark datasets from the annual Conference on Machine Translation (WMT).

The analysis reveals several critical findings. First, modern neural quality estimation metrics vastly outperform traditional string-matching metrics; CometKiwiQE22 achieved the highest overall pairwise accuracy with human judgment at 81.5%, closely followed by xCOMETXXL at 81.4%, while BLEU scored only 70.3%. Second, lexical string-matching metrics like BLEU never reach high levels of human agreement; while CometKiwiQE22 achieves 90% human agreement at a score gain of roughly 0.85 points, BLEU never reaches a 90% human agreement threshold even with large point gains. Third, string-based metrics fail completely when comparing unrelated translation systems developed by different teams; a 2-point BLEU gain yields approximately 90% agreement on iterated models but drops to roughly 55%—barely better than a coin toss—on unrelated models. Finally, the study demonstrated that metric delta accuracy remains stable across varying test set sizes, whereas traditional statistical significance testing (p-values) artificially deflates simply as test set size increases, risking false confidence in imperceptible gains.

These findings have direct operational and financial implications for machine translation workflows. Relying on legacy metrics or misinterpreting modern neural scores risks deploying models with no perceptible user benefit or rejecting genuine improvements. Furthermore, because p-values become arbitrarily small on large test sets, relying solely on statistical significance tests can lead teams to over-invest in models that do not deliver real-world quality gains.

The article recommends adopting CometKiwiQE22 as the primary evaluation metric because it avoids human reference bias while maintaining top-tier accuracy. Practitioners should pair it with a structurally distinct reference-based metric, such as BLEURT20, and report estimated human accuracy alongside raw score differences. In contrast, teams should strictly avoid using string-based metrics like BLEU or ChrF when comparing unrelated architectures, and avoid evaluating models with the same metrics used during system training or decoding.

These findings are supported with high confidence by an unprecedented volume of human judgment data across diverse domains. However, users should note that the underlying evaluation dataset primarily includes traditional translation architectures rather than large language model (LLM) translators, and human evaluators themselves exhibit natural noise and subjectivity. As a result, automated metric thresholds provide strong guidance for rapid assessment but cannot entirely replace direct human evaluation for high-stakes deployment decisions.

arXiv: 2401.06760kocmitom/MT-Thresholds
Cover for Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies

Abstract

Ten years ago, a single metric, BLEU, governed progress in machine translation research. For better or worse, there is no such consensus today, and consequently it is difficult for researchers to develop and retain intuitions about metric deltas that drove earlier research and deployment decisions. This paper investigates the “dynamic range” of a number of modern metrics in an effort to provide a collective understanding of the meaning of differences in scores both within and among metrics; in other words, we ask what point difference x in metric y is required between two systems for humans to notice? We conduct our evaluation on a new large dataset, ToShip23, using it to discover deltas at which metrics achieve system-level differences that are meaningful to humans, which we measure by pairwise system accuracy. We additionally show that this method of establishing delta-accuracy is more stable than the standard use of statistical p-values in regards to testset size. Where data size permits, we also explore the effect of metric deltas and accuracy across finer-grained features such as translation direction, domain, and system closeness.

Table of Contents

  • 1 Introduction
  • 2 Experimental Setup
  • 3 Unifying Metric Ranges
  • 3.1 Various Ranges for Metric Deltas
  • 3.2 Accuracy of Metric Deltas
  • 3.3 Aligning Metrics on Accuracy
  • 4 Factors Affecting Metric Deltas
  • 4.1 Different Domains and Datasets
  • 4.2 Language Pair
  • 4.3 Iterated versus Unrelated Systems
  • 4.4 Testset Size
  • 5 Discussion
  • 5.1 Best-performing Metrics
  • 5.2 Recommendations for MT Evaluation
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Metric Implementation Details
  • B ToShip23 Dataset Details
  • C Metrics Disagreement on Ranking
  • D Number of System Pairs in a Bin

Knowls

  1. Knowl 1 — Metric Delta Thresholds for Human Pairwise Accuracy Levels

    data/table

    Different machine translation evaluation metrics operate over vastly different score ranges (dynamic ranges) and exhibit different degrees of correlation with human judgments. Consequently, an identical numerical score difference between two translation systems (a metric delta Δ\Delta) represents different likelihoods that human annotators will agree with the ranking.

    By binning pairwise system deltas on the ToShip23 dataset (6,530 system pairs across 94 languages) and fitting parametrized sigmoid curves to empirical pairwise accuracy against human judgments, the minimum score delta Δ\Delta required to achieve a target pairwise human agreement level is established for each metric:

    Metric 50% 55% 60% 65% 70% 75% 80% 85% 90% 95%
    BLEU 0.27 0.52 0.78 1.06 1.39 1.79 2.34 3.35 – –
    ChrF 0.14 0.33 0.54 0.76 1.00 1.28 1.63 2.12 3.05 –
    spBLEU200 0.25 0.52 0.82 1.13 1.49 1.91 2.46 3.28 5.57 –
    Bleurtdefault 0.23 0.66 1.11 1.59 2.11 2.71 3.43 4.39 5.98 –
    Bleurt20 0.02 0.17 0.33 0.49 0.66 0.85 1.07 1.35 1.73 2.44
    Comet20 0.08 0.36 0.65 0.96 1.29 1.67 2.10 2.66 3.45 5.10
    Comet22 0.03 0.10 0.18 0.26 0.35 0.45 0.56 0.71 0.94 1.53
    Comet21QE 0.003 0.008 0.013 0.019 0.025 0.032 0.041 0.052 0.073 –
    CometKiwi22QE 0.01 0.08 0.16 0.24 0.33 0.42 0.53 0.67 0.85 1.18
    xCOMETXXL 0.02 0.19 0.37 0.56 0.76 0.98 1.24 1.55 1.99 2.74

    Key observations:

    • BLEU, ChrF, spBLEU200, Bleurt\textsubscript{default}, and Comet\textsubscript{21}\textsuperscript{QE} never achieve 95% agreement with human rankings on the dataset. In fact, BLEU never reaches 90% agreement even for deltas exceeding 6.0 BLEU points.
    • A delta of +1.06 BLEU yields only 65% estimated accuracy, which is equivalent in expected human agreement to a +0.24 delta in CometKiwi\textsubscript{22}\textsuperscript{QE}.
    • Reaching 90% pairwise agreement requires Δ≥0.85\Delta \ge 0.85 for CometKiwi\textsubscript{22}\textsuperscript{QE}, Δ≥0.94\Delta \ge 0.94 for Comet22, Δ≥1.73\Delta \ge 1.73 for Bleurt20, and Δ≥3.05\Delta \ge 3.05 for ChrF.
  2. Knowl 2 — Parametrized Sigmoid Mapping from Metric Score Deltas to Estimated Pairwise Accuracy

    model/method

    To map raw system-level metric score deltas into a unified scale of estimated human pairwise accuracy, an empirical binning and curve-fitting procedure is used.

    For a given metric, system pairs are sorted by absolute system-level delta x=∣score(SysA)−score(SysB)∣x = |\text{score}(\text{Sys}_A) - \text{score}(\text{Sys}_B)|. For each evaluated delta level, a fixed-size bin of the 300 closest system pairs is formed to compute empirical pairwise accuracy against human judgments. Because empirical accuracy curves fluctuate, a bounded, two-parameter sigmoid function is fitted to the binned data using the Levenberg-Marquardt damped least-squares algorithm:

    f(x)=ϕ11+exp⁡(−ϕ2⋅x)f(x) = \frac{\phi_1}{1 + \exp(-\phi_2 \cdot x)}

    where x≥0x \ge 0 is the metric delta Δ\Delta, ϕ1∈(0,100]\phi_1 \in (0, 100] represents the asymptotic reliability ceiling (the maximum expected percentage agreement of the metric with human judgments), and ϕ2>0\phi_2 > 0 determines the rate of accuracy growth per unit increase in score delta.

    Fitted parameter values obtained on the ToShip23 dataset:

    • BLEU: ϕ1=88.3\phi_1 = 88.3, ϕ2=1.0\phi_2 = 1.0
    • ChrF: ϕ1=93.0\phi_1 = 93.0, ϕ2=1.1\phi_2 = 1.1
    • spBLEU200: ϕ1=91.0\phi_1 = 91.0, ϕ2=0.8\phi_2 = 0.8
    • Bleurt\textsubscript{default}: ϕ1=94.7\phi_1 = 94.7, ϕ2=0.5\phi_2 = 0.5
    • Bleurt20: ϕ1=98.3\phi_1 = 98.3, ϕ2=1.4\phi_2 = 1.4
    • Comet20: ϕ1=97.3\phi_1 = 97.3, ϕ2=0.7\phi_2 = 0.7
    • Comet22: ϕ1=96.2\phi_1 = 96.2, ϕ2=2.8\phi_2 = 2.8
    • Comet\textsubscript{21}\textsuperscript{QE}: ϕ1=93.8\phi_1 = 93.8, ϕ2=43.4\phi_2 = 43.4
    • CometKiwi\textsubscript{22}\textsuperscript{QE}: ϕ1=98.9\phi_1 = 98.9, ϕ2=2.7\phi_2 = 2.7
    • xCOMET\textsubscript{XXL}: ϕ1=98.9\phi_1 = 98.9, ϕ2=1.2\phi_2 = 1.2
  3. Knowl 3 — Degradation of Lexical Metric Accuracy When Evaluating Unrelated MT Systems

    empirical result

    The accuracy of automatic machine translation metrics depends heavily on whether the evaluated system pair consists of iterated systems (a baseline system compared against an incremental improvement produced by the same development team) or unrelated systems (distinct architectures or systems developed independently by different teams, as in WMT evaluations).

    Empirical evaluation on the ToShip23 dataset (1,607 iterated system pairs and 1,422 unrelated system pairs) reveals a severe failure mode for string-based matching metrics:

    • On iterated systems, a +2.0 BLEU delta corresponds to approximately 90% pairwise ranking accuracy with human judgments.
    • On unrelated systems, the identical +2.0 BLEU delta yields only ~55% accuracy, which is scarcely better than a random coin toss (50%).
    • Similar degradation on unrelated systems occurs for ChrF, spBLEU200, and Bleurt\textsubscript{default}.
    • In contrast, pretrained neural metrics such as CometKiwi\textsubscript{22}\textsuperscript{QE}, Comet22, Bleurt20, and xCOMET\textsubscript{XXL} maintain high ranking accuracy across both iterated and unrelated systems (e.g., CometKiwi\textsubscript{22}\textsuperscript{QE} reaches >85% accuracy on unrelated systems at Δ≥1.0\Delta \ge 1.0).

    Consequently, surface and lexical metrics (BLEU, ChrF, spBLEU) are unsuitable for comparing unrelated machine translation systems.

  4. Knowl 4 — Sample-Size Invariance of Metric Deltas versus Sample-Size Sensitivity of P-Values

    empirical result

    The reliability of statistical significance testing (pp-values) versus metric score deltas (Δ\Delta) behaves differently as parallel testset sizes vary.

    Sampling experiments on WMT23 system pairs across sentence counts ranging from under 500 to over 8,000 sentences demonstrate:

    1. Metric delta stability: The mean system-level metric delta Δ\Delta (e.g., in CometKiwi\textsubscript{22}\textsuperscript{QE}) remains largely constant regardless of testset size. The variance of Δ\Delta is elevated only for small testsets below 500 segments and stabilizes rapidly as sample size grows.
    2. P-value deflation: The pp-value from paired Student's tt-tests systematically decreases toward zero as testset sentence count increases, even when the underlying metric delta is imperceptibly small. This occurs because statistical power increases with sample size NN, allowing arbitrarily small, humanly imperceptible differences to achieve statistical significance (p<0.05p < 0.05 or p<0.01p < 0.01) provided the sample is large enough.

    Therefore, reporting statistical significance alone is insufficient to demonstrate meaningful MT system improvements. Metric deltas and estimated pairwise accuracy provide a sample-size-stable measure of humanly noticeable effect size, while significance testing primarily serves to rule out noise at small score deltas.

  5. Knowl 5 — System-Level Pairwise Accuracy Benchmark across Machine Translation Metrics

    data/table

    System-level pairwise ranking accuracy measures the fraction of system pairs where an automatic metric's ranking matches human evaluation judgments.

    Evaluation on the ToShip23 dataset (6,530 system pairs across 94 languages, partitioned into 2019--2021 and 2022--2023 evaluation subsets) and the WMT23 MQM human evaluation subset (249 system pairs) produces the following rankings:

    Metric ToShip23 (All) ToShip23 (22–23) ToShip23 (19–21) WMT23 (MQM)
    System Pairs (NN) 6530 1843 4687 249
    CometKiwi22QE 81.5 74.5 84.3 90.0
    xCOMETXXL 81.4 75.3 83.9 92.8
    Comet20 80.1 73.2 82.9 86.3
    Bleurt20 78.6 69.8 82.1 89.2
    Comet22 78.6 71.1 81.5 84.7
    Comet21QE 76.8 71.2 79.0 69.5
    ChrF 71.9 61.4 76.0 79.5
    spBLEU200 71.6 61.0 75.7 81.9
    BLEU 70.3 61.3 73.9 81.5
    Bleurtdefault 69.9 61.0 73.4 85.1

    Key conclusions:

    • The reference-free Quality Estimation (QE) metric CometKiwi\textsubscript{22}\textsuperscript{QE} achieves the highest overall accuracy on ToShip23 (81.5%), demonstrating that state-of-the-art QE metrics match or exceed reference-based metrics while avoiding human reference bias.
    • xCOMET\textsubscript{XXL} is the top-performing reference-based metric on both ToShip23 (81.4%) and WMT23 MQM (92.8%).
    • Lexical string-matching metrics (ChrF, spBLEU200, BLEU) and the lightweight default BLEURT-Tiny model (Bleurt\textsubscript{default}) perform substantially worse than all modern neural metrics.
  6. Knowl 6 — Pairwise System Ranking Accuracy Metric

    equation

    The alignment between an automatic evaluation metric and human judgment over a set of system pairs is evaluated using pairwise accuracy:

    Acc=∣{(A,B)∈P:sign(metricΔ(A,B))=sign(humanΔ(A,B))}∣∣P∣\text{Acc} = \frac{|\{ (A, B) \in \mathcal{P} : \text{sign}(\text{metric}\Delta(A, B)) = \text{sign}(\text{human}\Delta(A, B)) \}|}{|\mathcal{P}|}

    where:

    • P\mathcal{P} is the set of all evaluated system pairs (A,B)(A, B);
    • metricΔ(A,B)=scoremetric(A)−scoremetric(B)\text{metric}\Delta(A, B) = \text{score}_{\text{metric}}(A) - \text{score}_{\text{metric}}(B) is the system-level score difference assigned by the automatic metric between system AA and system BB;
    • humanΔ(A,B)=scorehuman(A)−scorehuman(B)\text{human}\Delta(A, B) = \text{score}_{\text{human}}(A) - \text{score}_{\text{human}}(B) is the corresponding score difference assigned by human annotators;
    • sign(z)\text{sign}(z) returns +1+1 if z>0z > 0, −1-1 if z<0z < 0, and 00 if z=0z = 0.

    A pair is considered correctly classified if both the automatic metric and human evaluators agree on which system is superior.

  7. Knowl 7 — The ToShip23 Machine Translation Evaluation Benchmark

    experimental setup

    ToShip23 is an evaluation dataset created to benchmark machine translation metrics against human judgments at scale. It extends the earlier ToShip21 dataset by incorporating modern neural and LLM-based MT systems from 2022 and 2023, expanding domain coverage, and upgrading human annotation protocols.

    Dataset specifications:

    • Size: 3,016,000 human-annotated segments across 6,752 translation systems, forming 6,530 unique pairwise system comparisons.
    • Language Coverage: 94 languages (restricted to languages supported by BERT and XLM-RoBERTa embeddings).
    • Domain Coverage: Over 10 distinct domains (expanded from the 2 domains, news and speech, in ToShip21).
    • Human Evaluation Protocols: Direct Assessment combined with Scalar Quality Metric (DA+SQM) and Multidimensional Quality Metrics (MQM), replacing older source-based Direct Assessment protocols.
    • Translation Direction: Authentic direction parallel testsets (original human source translated into human reference target, avoiding synthetic or reverse-translated testsets whenever possible).
  8. Knowl 8 — Cross-Domain Transfer and Translation Direction Robustness of Metric Accuracy Thresholds

    empirical result

    The delta-accuracy mapping derived on ToShip23 was tested for generalizability across independent datasets, translation directions, and language typologies:

    1. Validation on WMT22/WMT23: When applying ToShip23-derived accuracy thresholds to 1,414 system pairs from WMT22 and WMT23, the predicted estimated accuracy tracks the real empirical accuracy closely across most metrics. Comet22 exhibits slight underestimation (real WMT accuracy is higher than estimated), while Comet\textsubscript{21}\textsuperscript{QE} overestimates performance on WMT.
    2. Translation Direction: Partitioning ToShip23 into into-English (3,178 system pairs) and out-of-English (3,217 system pairs) shows comparable delta-accuracy curves, indicating that translation direction does not substantially distort the relationship between score delta and human agreement.
    3. Non-Latin Language Families (CJK): On Chinese, Japanese, and Korean system pairs (992 pairs), character/string-based metrics such as ChrF show noticeable performance deviations and lower accuracy for given deltas, whereas pretrained neural metrics (such as CometKiwi\textsubscript{22}\textsuperscript{QE}) remain robust across language families.
  9. Knowl 9 — Methodological Guidelines and Limitations for Metric Delta Interpretation in Machine Translation

    limitation

    While estimated delta-accuracies establish grounded thresholds for interpreting score improvements, their application is subject to methodological constraints:

    1. Primary Metric Selection: CometKiwi\textsubscript{22}\textsuperscript{QE} is recommended as the primary MT metric due to top empirical accuracy and immunity to reference bias. Evaluations should include at least one structurally distinct complementary metric (e.g., Bleurt20, which is reference-based and uses a non-COMET architecture).
    2. Self-Evaluation Bias: Evaluating systems with the same metric or backbone model used during training, Minimum Bayes Risk (MBR) decoding, or corpus filtering artificially inflates scores. Similarly, LLM-based evaluators favor translations generated by their own model family.
    3. Empirical Nature of Thresholds: Estimated delta-accuracies are derived from human annotations, which contain inherent rater noise, particularly on systems with very close performance. LLM-based translation models are underrepresented in the benchmark.
    4. Non-Rejection Principle: Estimated delta-accuracy should not be used as a hard cutoff to reject experimental findings, just as low significance pp-values should not be used as sole rejection criteria. Human evaluation remains necessary to confirm definitive quality improvements.

Coverage note — None; all core methodological steps, empirical findings, datasets, thresholds, and evaluation analyses from the paper's contribution are captured across the knowls.

References

  1. 1.Chantal Amrhein, Nikita Moghe, and Liane Guillou. 2022. ACES: Translation accuracy challenge sets for evaluating machine translation metrics. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 479–513, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Anja Belz and Ehud Reiter. 2006. Comparing automatic and human evaluation of NLG systems. In 11th Conference of the European Chapter of the Association for Computational Linguistics, pages 313–320, Trento, Italy. Association for Computational Linguistics.
  4. 4.Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein. 2012. An empirical investigation of statistical significance in NLP. In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, pages 995–1005, Jeju Island, Korea. Association for Computational Linguistics.
  5. 5.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 131–198, Berlin, Germany. Association for Computational Linguistics.
  6. 6.Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006. Re-evaluating the role of Bleu in machine translation research. In 11th Conference of the European Chapter of the Association for Computational Linguistics, pages 249–256, Trento, Italy. Association for Computational Linguistics.
  7. 7.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  8. 8.Etienne Denoual and Yves Lepage. 2005. BLEU in characters: Towards automatic MT evaluation in languages without word delimiters. In Companion Volume to the Proceedings of Conference including Posters/Demos and tutorial abstracts.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat. 2023. The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation. In Proceedings of the Eighth Conference on Machine Translation, pages 1066–1083, Singapore. Association for Computational Linguistics.
  11. 11.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  12. 12.Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022a. High quality rather than high model probability: Minimum Bayes risk decoding with neural metrics. Transactions of the Association for Computational Linguistics, 10:811–825.
  13. 13.Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent. In Proceedings of the Eighth Conference on Machine Translation, pages 578–628, Singapore. Association for Computational Linguistics.
  14. 14.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022b. Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 46–68, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  15. 15.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  16. 16.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  17. 17.Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman. 2016. Statistical tests, p values, confidence intervals, and power: A guide to misinterpretations. European journal of epidemiology, 31:337–350.
  18. 18.Eduard Hovy and Deepak Ravichandran. 2003. Holy and unholy grails. In Proceedings of Machine Translation Summit IX: Plenaries, New Orleans, USA.
  19. 19.Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022. DEMETR: Diagnosing evaluation metrics for translation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9540–9561, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  20. 20.Ken Kelley and Kristopher J Preacher. 2012. On effect size. Psychological methods, 17(2):137.
  21. 21.Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022. Findings of the 2022 conference on machine translation (WMT22). In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 1–45, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  22. 22.Tom Kocmi and Christian Federmann. 2023. GEMBA-MQM: Detecting translation quality error spans with GPT-4. In Proceedings of the Eighth Conference on Machine Translation, pages 768–775, Singapore. Association for Computational Linguistics.
  23. 23.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation. In Proceedings of the Sixth Conference on Machine Translation, pages 478–494, Online. Association for Computational Linguistics.
  24. 24.Kenneth Levenberg. 1944. A method for the solution of certain non-linear problems in least squares. Quarterly of Applied Mathematics.
  25. 25.Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023. Llms as narcissistic evaluators: When ego inflates evaluation scores. arXiv preprint arXiv:2311.09766.
  26. 26.Chi-kiu Lo, Rebecca Knowles, and Cyril Goutte. 2023. Beyond correlation: Making sense of the score differences of new MT evaluation metrics. In Proceedings of Machine Translation Summit XIX, Vol. 1: Research Track, pages 186–199.
  27. 27.Benjamin Marie. 2022. Yes, we need statistical significance testing.
  28. 28.Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021. Scientific credibility of machine translation research: A meta-evaluation of 769 papers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7297–7306, Online. Association for Computational Linguistics.
  29. 29.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020a. Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4984–4997, Online. Association for Computational Linguistics.
  30. 30.Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020b. Results of the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 688–725, Online. Association for Computational Linguistics.
  31. 31.Allan Mori, Gustavo Vale, Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo, Elder Cirilo, Pooyan Jamshidi, and Christian Kastner. 2018. Evaluating domain-specific metric thresholds: An empirical study. In Proceedings of the 2018 International Conference on Technical Debt, pages 41–50.
  32. 32.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  33. 33.Jan-Thorsten Peter, David Vilar, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, and Markus Freitag. 2023. There’s no data like better data: Using QE metrics for MT data filtering. In Proceedings of the Eighth Conference on Machine Translation, pages 561–577, Singapore. Association for Computational Linguistics.
  34. 34.Luke Plonsky and Frederick L Oswald. 2014. How big is “big”? interpreting effect sizes in l2 research. Language Learning.
  35. 35.Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  36. 36.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021. Learning compact metrics for MT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 751–762, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  38. 38.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  39. 39.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  40. 40.Raed Shatnawi, Wei Li, James Swain, and Tim Newman. 2010. Finding software metrics threshold values using ROC curves. J. Softw. Maint. Evol., 22(1):1–16.
  41. 41.Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018. Attaining the unattainable? reassessing claims of human parity in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 113–123, Brussels, Belgium. Association for Computational Linguistics.
  42. 42.Vilém Zouhar and Ondřej Bojar. 2024. Quality and quantity of machine translation references for automated metrics.

Citation

MLA
Kocmi, T., et al. “Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1999–2014, https://doi.org/10.18653/v1/2024.acl-long.110.
APA
Kocmi, T., Zouhar, V., Federmann, C., & Post, M. (2024). Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1999–2014. https://doi.org/10.18653/v1/2024.acl-long.110
Chicago
Kocmi, T., V. Zouhar, C. Federmann, and M. Post. 2024. “Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1999–2014. https://doi.org/10.18653/v1/2024.acl-long.110.
Harvard
Kocmi, T. et al. (2024) “Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1999–2014. Available at: https://doi.org/10.18653/v1/2024.acl-long.110.
Vancouver
1. Kocmi T, Zouhar V, Federmann C, Post M (2024) Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1999–2014

BibTeX

@inproceedings{kocmi-etal-2024-navigating,
    title = "Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies",
    author = "Kocmi, Tom  and
      Zouhar, Vil{\'e}m  and
      Federmann, Christian  and
      Post, Matt",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.110/",
    doi = "10.18653/v1/2024.acl-long.110",
    pages = "1999--2014"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/