Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Daniel DeutschGeorge F. FosterMarkus Freitag

article2023EMNLP92 citations

Proposes a tie-aware pairwise accuracy metric and a tie calibration procedure to fix vulnerabilities in standard Kendall's tau variants, preventing evaluation gaming and ensuring fair ranking assessments for machine translation metrics.

Listen

Evaluating the performance of automatic evaluation metrics for machine translation and generative AI—a process called meta-evaluation—is critical for selecting the best translation systems and guiding automated decoding methods. Traditionally, meta-evaluation at the segment level relies on ranking statistics such as variants of Kendall's tau to measure how well metric rankings correlate with human judgments. However, as translation systems improve, expert human evaluators increasingly encounter identical, error-free outputs, resulting in a large proportion of tied scores (often 40% to 53% of all evaluated pairs). At the same time, newer metrics based on large language models and discrete error tagging frequently output ties. Existing ranking statistics do not properly handle these ties, introducing major blind spots and distortion into metric benchmarks.

The main objective of the article is to demonstrate that conventional ranking correlation metrics are vulnerable to severe biases and potential gaming when ties are present, and to establish a fairer, more interpretable evaluation framework using pairwise accuracy paired with a tie calibration procedure.

The authors conducted statistical and empirical analyses using human expert Multidimensional Quality Metrics data from the WMT'22 shared task across three language pairs (English–German, Chinese–English, and English–Russian), covering 13 to 15 translation systems and tens of thousands of segment pairs. They evaluated standard regression metrics (such as Metric-X and COMET-22) alongside discrete and large language model metrics (such as MaTESe and GEMBA). They examined how different correlation variants handle ties, tested synthetic score-bucketing strategies to assess metric gaming, and implemented an exact algorithm to find optimal pairwise tie thresholds.

The analysis produced several key findings. First, existing variants of Kendall's tau fail significantly in the presence of ties: some heavily penalize metrics for predicting ties by treating them as errors (dropping valid metrics to the bottom of leaderboards), while others ignore constant scores and produce undefined ("NaN") outputs. Second, this undefined-output issue can be gamed; bucketing continuous metric scores to artificially trigger ties on difficult translation groups reduced the evaluated segment count but falsely boosted correlation scores by large margins. Third, replacing Kendall's tau with pairwise accuracy directly credits metrics for correctly predicting ties while eliminating undefined evaluation groups. Finally, applying the proposed tie calibration algorithm—which automatically identifies an optimal score difference threshold below which pairs are treated as ties—restored continuous regression metrics like Metric-X and COMET to top rankings while enabling fair, direct comparison with discrete classification-based metrics.

These findings indicate that existing benchmark rankings have mischaracterized metric quality by either artificially favoring or punishing models based purely on how they format score ties rather than true accuracy. In practical deployments, relying on flawed ranking statistics creates substantial risk of adopting sub-optimal metrics for production pipelines, automated system tuning, and quality assurance. Moving to an intuitive pairwise accuracy metric—defined simply as the proportion of pairs correctly ordered or correctly tied—provides organizations with transparent, interpretable evaluations.

To ensure fair and reliable meta-evaluations, organizations and benchmark organizers should adopt pairwise accuracy combined with tie calibration for ranking-based metric assessments. Developers can also use the calibrated threshold as an operational indicator of when a metric score difference is practically meaningful. However, practitioners must note that calibrated tie thresholds currently do not generalize well across datasets with vastly different tie distributions (such as across different years or language pairs) and reflect pair-level decisions rather than a global transitivity ranking. While confidence in the mathematical robustness and empirical stability of pairwise accuracy is high, benchmarking across substantially new domains will require dataset-specific calibration until metrics natively predict ties more accurately.

Deutsch et al (2023).pdf
Cover for Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Abstract

Kendall’s τ is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the literature. We demonstrate that, in settings like modern MT meta-evaluation, existing variants have weaknesses arising from their handling of ties, and in some situations can even be gamed. We propose instead to meta-evaluate metrics with a version of pairwise accuracy that gives metrics credit for correctly predicting ties, in combination with a tie calibration procedure that automatically introduces ties into metric scores, enabling fair comparison between metrics that do and do not predict ties. We argue and provide experimental evidence that these modifications lead to fairer ranking-based assessments of metric performance.

Table of Contents

  • 1 Introduction
  • 2 Background & Related Work
  • 2.1 Why not Pearson or Spearman?
  • 2.2 Metric Meta-Evaluation
  • 2.3 The Landscape of Kendall's τ
  • 3 Analysis Setup
  • 4 Why Ties are Important
  • 5 Shortcomings of Kendall's Variants
  • 5.1 A Motivating Example
  • 5.2 The NaN Problem
  • 6 Evaluating with Pairwise Accuracy
  • 6.1 Evaluating Ties and Non-Ties
  • 7 Tie Calibration
  • 8 Analysis
  • 8.1 Comparing Metric Rankings
  • 8.2 Generalization of Epsilon
  • 8.3 Where are Ties Introduced?
  • 8.4 Class-Specific Statistics
  • 9 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Correlation Definitions
  • B WMTTabular Notation
  • C Additional Results
  • D Tie Calibration Pseudocode
  • E Epsilon Search Approximation
  • F Unbabel Normalization

Knowls

  1. Knowl 1 — Tie-aware pairwise accuracy

    definition

    For nn translations with human scores hih_i and metric scores mim_i, evaluate all (n2)\binom{n}{2} unordered pairs. A pair is concordant (CC) when the human and metric scores order it the same way, discordant (DD) when they order it oppositely, tied only in the human scores (ThT_h), tied only in the metric scores (TmT_m), or tied in both (ThmT_{hm}). The proposed tie-aware pairwise accuracy is

    acc⁡eq=C+ThmC+D+Th+Tm+Thm.\operatorname{acc}_{eq}=\frac{C+T_{hm}}{C+D+T_h+T_m+T_{hm}}.

    It is the fraction of pairs whose ordering is correct or for which the metric correctly predicts a tie. Thus, unlike the Kendall variants examined in the paper, it gives direct credit for correctly matching human ties and always yields a defined value when there is at least one pair. Its range is [0,1][0,1].

  2. Knowl 2 — Tie calibration selects a score-difference threshold

    algorithm

    Tie calibration makes metrics that output few exact ties comparable with metrics that output many. Given human scores hih_i, metric scores mim_i, and a tie-sensitive objective such as acc⁡eq\operatorname{acc}_{eq}, it selects an absolute score-difference threshold ϵ\epsilon and treats a pair as tied by the metric whenever ∣mi−mj∣≤ϵ|m_i-m_j|\leq\epsilon. The selected threshold maximizes the chosen objective. For grouped correlations, the same threshold is evaluated using the specified aggregation over groups, with pair counts maintained separately for each group.

    An exact search considers every distinct pairwise absolute score difference as a candidate threshold. Starting at ϵ=0\epsilon=0, it sorts pairs by their score differences and sweeps upward through those candidates, updating pair classifications and the objective as each batch of equally distant pairs becomes tied. It returns a maximizing threshold and its objective value; ties between equally good thresholds need not be resolved in a particular way.

    Input: Human scores h, metric scores m, objective Q, and optionally group labels
    Output: A maximizing threshold epsilon_best and its objective value
    Create one record for each pair (i,j) with i < j, storing d = abs(m_i - m_j)
    Initialize pair counts by treating a metric pair as tied exactly when d = 0
    Set epsilon_best = 0 and best_value = Q(pair counts)
    Sort pair records by d in ascending order
    For each distinct positive value d in sorted order:
        For every pair record with this value d:
            Remove the pair from its current category
            If h_i = h_j, add it to the category tied in both human and metric scores
            Otherwise, add it to the category tied only in metric scores
            If groups are used, update the counts for this pair's group
        Evaluate Q after processing the entire batch with difference d
        If the value exceeds best_value, store d as epsilon_best and update best_value
    Return epsilon_best and best_value

    The exact search takes O(n2log⁡n)O(n^2\log n) time for nn scored translations, dominated by sorting the (n2)\binom{n}{2} pairs. For very large pair sets, the authors downsampled candidate pairs; in their approximation analysis, sampling 10% of the pairs produced close results over 30 runs, with maximum observed differences of 2.5×10−32.5\times10^{-3} in ϵ∗\epsilon^* and 4.3×10−54.3\times10^{-5} in accuracy.

  3. Knowl 3 — Existing Kendall variants miss correct-tie information

    empirical result

    A six-item example illustrates that Kendall variants can fail to reflect whether a metric predicts human ties. Let the human scores be h=[0,0,0,0,1,2]h=[0,0,0,0,1,2], and compare m1=[0,0,0,0,2,1]m_1=[0,0,0,0,2,1] with m2=[0,1,2,3,4,5]m_2=[0,1,2,3,4,5]. Of the 15 possible pairs, m1m_1 orders only one incorrectly, while m2m_2 orders six incorrectly. Yet the existing variants give no direct credit for the four-item tied block that m1m_1 predicts correctly: for example, τ10\tau_{10} is 0.780.78 for m1m_1 but 1.01.0 for m2m_2, and τ13\tau_{13} and τ14\tau_{14} also give m2m_2 a value of 1.01.0. By contrast, tie-aware pairwise accuracy is 0.930.93 for m1m_1 and 0.600.60 for m2m_2. The example shows why a statistic that does not reward correct tie predictions can rank a substantially worse pairwise predictor too highly.

  4. Knowl 4 — Undefined grouped correlations can be exploited

    empirical result

    When a human-score or metric-score vector within a group is constant, many Kendall correlations are undefined. In grouped segment-level evaluation, undefined group correlations are omitted from the average in practice, so a metric can raise its reported correlation by creating ties in difficult groups and thereby removing those groups from evaluation. The authors demonstrated that equal-width bucketing of Metric-X scores can create this effect: the resulting correlations use fewer non-undefined groups, so scores across different bucketings are not comparable.

    In WMT’22 en→de group-by-item evaluation, MaTESe’s τb=0.281\tau_b=0.281 was calculated on 773 non-undefined groups, whereas Metric-X’s τb=0.270\tau_b=0.270 used 1,133 groups. On the same 773 groups, Metric-X reached 0.2960.296, exceeding MaTESe. Thus apparent metric rankings can reflect differing evaluation subsets rather than just metric quality. Pairwise accuracy avoids this particular constant-vector problem for groups containing at least one pair because its denominator includes every pair and is nonzero.

  5. Knowl 5 — WMT’22 data contain many human ties, while metric tie rates vary sharply

    experimental setup

    The main analysis uses expert MQM ratings from the WMT’22 metrics shared task for en→de, zh→en, and en→ru. Each language pair has 13–15 systems and roughly 1,300–1,900 rated segments per system. In group-by-item evaluation, human-score ties are common, and many are ties at MQM score zero, indicating error-free translations. The WMT’22 counts reported on page 4 are:

    • en→de: 18k segments, 120k pairs, 64k tied human-score pairs (53%), including 48k zero-score pairs (40%).
    • zh→en: 28k segments, 197k pairs, 82k tied pairs (42%), including 63k zero-score pairs (32%).
    • en→ru: 20k segments, 138k pairs, 61k tied pairs (44%), including 40k zero-score pairs (29%).

    The percentage of group-by-item pairs tied by each metric also differed substantially. Metric-X’s rates were 0.7%, 0.2%, and 0.5% for en→de, zh→en, and en→ru; COMET-22’s were 1.3%, 0.1%, and 0.5%. In the same language-pair order, MaTESe tied 71.9%, 39.6%, and 80.8% of pairs; GEMBA-GPT-3.5 tied 60.3%, 56.6%, and 50.9%; and GEMBA-GPT-4 tied 69.6%, 46.9%, and 60.1%. These differences motivate a method that can compare metrics that predict exact ties at very different rates.

  6. Knowl 6 — Metric rankings change substantially under tie-aware evaluation

    empirical result

    On WMT’22 en→de, using group-by-item segment-level correlations, the choice of statistic changed metric rankings. The page 7 comparison reports the following values, with ranks in parentheses; acc⁡eq∗\operatorname{acc}_{eq}^{*} means pairwise accuracy after tie calibration, and ϵ∗\epsilon^* is the selected threshold in each metric’s score units:

    • Metric-X: τb=0.270\tau_b=0.270 (4), τ10=0.381\tau_{10}=0.381 (1), acc⁡eq∗=0.605\operatorname{acc}_{eq}^{*}=0.605 (1), ϵ∗=0.04\epsilon^*=0.04.
    • UniTE: 0.2780.278 (3), 0.3220.322 (3), 0.5950.595 (2), 0.140.14.
    • COMET-22: 0.2580.258 (5), 0.3660.366 (2), 0.5940.594 (3), 0.110.11.
    • MaTESe: 0.2810.281 (2), −0.459-0.459 (16), 0.5820.582 (4), 0.000.00.
    • GEMBA-GPT-4: 0.3220.322 (1), −0.367-0.367 (15), 0.5730.573 (6), 4.004.00.
    • GEMBA-GPT-3.5: 0.2090.209 (8), −0.344-0.344 (14), 0.5450.545 (15), 15.0015.00.
    • A constant metric that predicts a tie for every pair: 0.0000.000 (17), −1.000-1.000 (18), 0.5340.534 (18), 0.000.00.

    The strong rank changes are concentrated among metrics that output many ties. In particular, τ10\tau_{10} penalizes such predictions, while calibrated pairwise accuracy puts the best regression-style metrics near the top and does not systematically penalize metrics for their original tie rate.

  7. Knowl 7 — Pairwise accuracy can be decomposed into tie and ranking performance

    definition

    For the same pair categories used in tie-aware pairwise accuracy, class-specific precision and recall distinguish tie prediction from ranking non-tied pairs. Let CC be concordant pairs, DD discordant pairs, ThT_h pairs tied only by human scores, TmT_m pairs tied only by metric scores, and ThmT_{hm} pairs tied by both. Then:

    • Tie precision is Thm/(Thm+Tm)T_{hm}/(T_{hm}+T_m), and tie recall is Thm/(Thm+Th)T_{hm}/(T_{hm}+T_h).
    • Correct-rank precision is C/(C+D+Th)C/(C+D+T_h): the share of pairs the metric ranks that it ranks correctly.
    • Correct-rank recall is C/(C+D+Tm)C/(C+D+T_m): the share of human-nontied pairs the metric orders correctly.

    These measures, and their class-specific F1F_1 scores, expose performance differences that a single aggregate accuracy can hide when tied and non-tied pairs are imbalanced.

  8. Knowl 8 — A calibrated tie threshold may not transfer across datasets

    empirical result

    The optimal absolute threshold ϵ∗\epsilon^* can depend on dataset properties, particularly the frequency of human ties. When calibrated on one WMT dataset and applied to another, the en→de threshold changed by 0.030.03, and held-out pairwise accuracy changed by a relative 2%, indicating relatively stable transfer in that comparison. The zh→en result was different: WMT’21 had 23% tied human-score pairs versus 41% in WMT’22, and the threshold selected on one dataset did not generalize well to the other. The authors therefore treat ϵ∗\epsilon^* as a latent variable calibrated on the evaluation dataset; its usefulness for fair comparison does not imply that it is a portable threshold for interpreting scores across dissimilar datasets.

  9. Knowl 9 — Calibration tends to introduce ties among high-scoring pairs

    empirical result

    For Metric-X on WMT’22 zh→en, the page 8 score-distribution analysis found that pairs tied by the selected threshold were skewed toward higher average metric scores compared with all pairs. Because Metric-X correlates relatively well with MQM scores, the authors suggest that many introduced ties correspond to high-quality, possibly error-free translations; this interpretation is a hypothesis rather than a demonstrated causal explanation. In a separate analysis of COMET-22 on WMT’22 en→de, tie-prediction F1F_1 was much higher than correct-rank F1F_1 across almost all tested thresholds, suggesting stronger performance on predicting ties than on ordering non-tied pairs. The authors attribute this pattern as likely related to the abundance of perfect translations and the high-score bias in where calibration introduces ties.

  10. Knowl 10 — Tie calibration has scale and transitivity limitations

    limitation

    Tie calibration assumes that an absolute difference in metric scores represents the same quality difference everywhere on a metric’s scale. For example, it treats differences of 0.10.1 at scores near 0.10.1 and near 100.1100.1 as equivalent. The authors tried relative score differences and found little change in their results, but note that other metrics may behave differently. In addition, pairwise thresholding does not necessarily produce a globally consistent ordering: with scores 1,2,31,2,3 and ϵ=1\epsilon=1, the first and second scores are tied, as are the second and third, while the first and third are not. Finally, the authors argue that their proposal is fairer based on its behavior and experiments, but report no way to prove that fairness claim.

Coverage note — The supplementary replication using Unbabel MQM normalization and the full per-language, per-group metric-ranking tables are omitted because they largely repeat the main analyses; the paper reports that the normalization change does not alter its central conclusions.

References

  1. 1.Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017. Results of the WMT17 Metrics Shared Task. In Proceedings of the Second Conference on Machine Translation, pages 489–513, Copenhagen, Denmark. Association for Computational Linguistics.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-Shot Learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Chris Callison-Burch, Philipp Koehn, Christof Monz, Kay Peterson, Mark Przybocki, and Omar Zaidan. 2010. Findings of the 2010 Joint Workshop on Statistical Machine Translation and Metrics for Machine Translation. In Proceedings of the Joint Fifth Workshop on Statistical Machine Translation and MetricsMATR, pages 17–53, Uppsala, Sweden. Association for Computational Linguistics.
  4. 4.Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022. Quality-Aware Decoding for Neural Machine Translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1396–1412, Seattle, United States. Association for Computational Linguistics.
  5. 5.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a. Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  6. 6.Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022a. High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics. Transactions of the Association for Computational Linguistics, 10:811–825.
  7. 7.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022b. Results of WMT22 Metrics Shared Task: Stop Using BLEU – Neural Metrics Are Better and More Robust. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 46–68, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  8. 8.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b. Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  9. 9.Maurice G Kendall. 1938. A New Measure of Rank Correlation. Biometrika, 30(1/2):81–93.
  10. 10.Maurice G Kendall. 1945. The Treatment of Ties in Ranking Problems. Biometrika, 33(3):239–251.
  11. 11.Tom Kocmi and Christian Federmann. 2023. Large Language Models Are State-of-the-Art Evaluators of Translation Quality.
  12. 12.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation. In Proceedings of the Sixth Conference on Machine Translation, pages 478–494, Online. Association for Computational Linguistics.
  13. 13.Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014. Multidimensional Quality Metrics (MQM): A Framework for Declaring and Describing Translation Quality Metrics. Tradumàtica, (12):0455–463.
  14. 14.Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018. Results of the WMT18 Metrics Shared Task: Both characters and embeddings achieve good performance. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 671–688, Belgium, Brussels. Association for Computational Linguistics.
  15. 15.Matouš Macháček and Ondřej Bojar. 2013. Results of the WMT13 Metrics Shared Task. In Proceedings of the Eighth Workshop on Statistical Machine Translation, pages 45–51, Sofia, Bulgaria. Association for Computational Linguistics.
  16. 16.Matouš Macháček and Ondřej Bojar. 2014. Results of the WMT14 Metrics Shared Task. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 293–301, Baltimore, Maryland, USA. Association for Computational Linguistics.
  17. 17.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020. Tangled up in bleu: Reevaluating the evaluation of automatic machine translation evaluation metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4984–4997.
  18. 18.Evgeny Matusov, Gregor Leusch, Oliver Bender, and Hermann Ney. 2005. Evaluating Machine Translation Output with Automatic Sentence Segmentation. In Proceedings of the Second International Workshop on Spoken Language Translation, Pittsburgh, Pennsylvania, USA.
  19. 19.Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Niccolò Campolungo, and Roberto Navigli. 2022. MaTESe: Machine Translation Evaluation as a Sequence Tagging Problem. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 569–577, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  20. 20.Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022. COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task. In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 578–585, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  21. 21.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A Neural Framework for MT Evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  22. 22.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  23. 23.Alan Stuart. 1953. The Estimation and Comparison of Strengths of Association in Contingency Tables. Biometrika, 40(1/2):105–110.

Citation

MLA
Deutsch, D., et al. “Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12914–29, https://doi.org/10.18653/v1/2023.emnlp-main.798.
APA
Deutsch, D., Foster, G., & Freitag, M. (2023). Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12914–12929. https://doi.org/10.18653/v1/2023.emnlp-main.798
Chicago
Deutsch, D., G. Foster, and M. Freitag. 2023. “Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12914–29. https://doi.org/10.18653/v1/2023.emnlp-main.798.
Harvard
Deutsch, D., Foster, G. and Freitag, M. (2023) “Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12914–12929. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.798.
Vancouver
1. Deutsch D, Foster G, Freitag M (2023) Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12914–12929

BibTeX

@inproceedings{deutsch-etal-2023-ties,
    title = "Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration",
    author = "Deutsch, Daniel  and
      Foster, George  and
      Freitag, Markus",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.798/",
    doi = "10.18653/v1/2023.emnlp-main.798",
    pages = "12914--12929"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/