Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Daniel DeutschGeorge F. FosterMarkus Freitag
Proposes a tie-aware pairwise accuracy metric and a tie calibration procedure to fix vulnerabilities in standard Kendall's tau variants, preventing evaluation gaming and ensuring fair ranking assessments for machine translation metrics.
Evaluating the performance of automatic evaluation metrics for machine translation and generative AI—a process called meta-evaluation—is critical for selecting the best translation systems and guiding automated decoding methods. Traditionally, meta-evaluation at the segment level relies on ranking statistics such as variants of Kendall's tau to measure how well metric rankings correlate with human judgments. However, as translation systems improve, expert human evaluators increasingly encounter identical, error-free outputs, resulting in a large proportion of tied scores (often 40% to 53% of all evaluated pairs). At the same time, newer metrics based on large language models and discrete error tagging frequently output ties. Existing ranking statistics do not properly handle these ties, introducing major blind spots and distortion into metric benchmarks.
The main objective of the article is to demonstrate that conventional ranking correlation metrics are vulnerable to severe biases and potential gaming when ties are present, and to establish a fairer, more interpretable evaluation framework using pairwise accuracy paired with a tie calibration procedure.
The authors conducted statistical and empirical analyses using human expert Multidimensional Quality Metrics data from the WMT'22 shared task across three language pairs (English–German, Chinese–English, and English–Russian), covering 13 to 15 translation systems and tens of thousands of segment pairs. They evaluated standard regression metrics (such as Metric-X and COMET-22) alongside discrete and large language model metrics (such as MaTESe and GEMBA). They examined how different correlation variants handle ties, tested synthetic score-bucketing strategies to assess metric gaming, and implemented an exact algorithm to find optimal pairwise tie thresholds.
The analysis produced several key findings. First, existing variants of Kendall's tau fail significantly in the presence of ties: some heavily penalize metrics for predicting ties by treating them as errors (dropping valid metrics to the bottom of leaderboards), while others ignore constant scores and produce undefined ("NaN") outputs. Second, this undefined-output issue can be gamed; bucketing continuous metric scores to artificially trigger ties on difficult translation groups reduced the evaluated segment count but falsely boosted correlation scores by large margins. Third, replacing Kendall's tau with pairwise accuracy directly credits metrics for correctly predicting ties while eliminating undefined evaluation groups. Finally, applying the proposed tie calibration algorithm—which automatically identifies an optimal score difference threshold below which pairs are treated as ties—restored continuous regression metrics like Metric-X and COMET to top rankings while enabling fair, direct comparison with discrete classification-based metrics.
These findings indicate that existing benchmark rankings have mischaracterized metric quality by either artificially favoring or punishing models based purely on how they format score ties rather than true accuracy. In practical deployments, relying on flawed ranking statistics creates substantial risk of adopting sub-optimal metrics for production pipelines, automated system tuning, and quality assurance. Moving to an intuitive pairwise accuracy metric—defined simply as the proportion of pairs correctly ordered or correctly tied—provides organizations with transparent, interpretable evaluations.
To ensure fair and reliable meta-evaluations, organizations and benchmark organizers should adopt pairwise accuracy combined with tie calibration for ranking-based metric assessments. Developers can also use the calibrated threshold as an operational indicator of when a metric score difference is practically meaningful. However, practitioners must note that calibrated tie thresholds currently do not generalize well across datasets with vastly different tie distributions (such as across different years or language pairs) and reflect pair-level decisions rather than a global transitivity ranking. While confidence in the mathematical robustness and empirical stability of pairwise accuracy is high, benchmarking across substantially new domains will require dataset-specific calibration until metrics natively predict ties more accurately.
- Paper: COMET: A Neural Framework for MT Evaluation, Ricardo Rei et al. (2020). The paper benchmarks COMET among modern translation metrics, so understanding this neural metric’s design clarifies what is being evaluated.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). The paper compares learned metrics such as BLEURT with alternative evaluators, and BLEURT’s metric design helps contextualize those benchmark results.
- Paper: Bleu: a Method for Automatic Evaluation of Machine Translation, Kishore Papineni et al. (2002). BLEU’s role as a conventional translation metric provides context for the paper’s contrast between established ranking statistics and newer metric types.
- Paper: Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies, Tom Kocmi et al. (2024). This later study carries forward pairwise accuracy as an interpretable measure, extending the paper’s focus on fair metric comparison to score magnitudes and meaningful quality differences.
