On the Limitations of Reference-Free Evaluations of Generated Text

Daniel DeutschRotem DrorDan Roth

article2022EMNLP71 citations

Demonstrates that reference-free text evaluation metrics act as generation models themselves, exposing critical flaws where metrics favor models similar to their own architecture and penalize superior human-written outputs.

Listen

Automated evaluation of text generation systems, such as machine translation and document summarization, traditionally relies on comparing outputs against human-written reference texts. Because gathering human references is costly, slow, and often infeasible for real-time applications, researchers have increasingly developed reference-free evaluation metrics that score candidate text directly from the input source. However, adopting reference-free metrics as primary benchmarks introduces fundamental risks of mismeasuring model performance.

The article demonstrates that reference-free evaluation metrics are mathematically equivalent to text generation models and inherently limited by significant structural biases. Specifically, it assesses whether reference-free metrics can be safely used to measure genuine task performance and track overall research progress.

The authors conducted empirical evaluations across machine translation and summarization using established benchmarks, including WMT19 data spanning 18 language pairs as well as the SummEval and REALSumm datasets covering 16 to 25 summarization systems. They analyzed three prominent reference-free metrics—Prism-src, COMET-QE, and QuestEval—by formulating simple search algorithms, such as beam search, greedy sentence selection, and output reranking, to directly optimize each metric at inference time without human references.

The analysis produced three critical findings. First, simple optimization procedures reliably generated outputs that outperformed standard baseline models on reference-free metrics, such as a 38% relative improvement in COMET-QE on German-to-English translation. Second, despite their top reference-free scores, these optimized outputs exhibited average or below-average actual quality when judged by standard reference-based metrics and contained severe translation errors. Third, reference-free metrics routinely scored machine outputs higher than human-written text. In German-to-English translation under Prism-src, nearly all automated models received higher scores than authentic human translations, showing a clear bias against genuine human language. Furthermore, scoring systems against a metric's optimized output showed an exceptionally strong correlation (average Pearson r of 0.88 to 0.95) with the reference-free scores, confirming that reference-free metrics effectively treat their own generated text as a gold standard.

These findings imply that using reference-free metrics to guide system development creates a misleading feedback loop. Instead of learning to generate high-quality, human-like text, models optimize toward the specific quirks and systemic errors of the evaluation metric’s internal model. Relying on these metrics as primary performance benchmarks risks misdirecting research investments, adopting lower-quality language systems, and penalizing genuinely superior models that diverge from the evaluator model's patterns.

The authors recommend that organizations stop using reference-free metrics as primary benchmarks or objective functions for measuring task progress. When measuring overall generation quality, stakeholders should invest in collecting human-written references for evaluation. Reference-free metrics should instead be repurposed as diagnostic tools or quality estimation safeguards to flag potential catastrophic errors, verify source faithfulness, or measure text fluency.

While the theoretical equivalence between reference-free metrics and generation models applies universally, the empirical demonstrations in the article are limited to machine translation and summarization datasets. The authors also relied on established reference-based metrics rather than expensive, ground-truth human annotations to evaluate the optimized outputs. Nonetheless, the evidence strongly warns against treating reference-free scores as substitutes for reference-based or human evaluation.

arXiv: 2210.12563
Cover for On the Limitations of Reference-Free Evaluations of Generated Text

Abstract

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications. However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization. We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans. Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.¹

Table of Contents

  • 1 Introduction
  • 2 Reference-Free Metrics as Models
  • 3 Analysis Setup
  • 4 Metric Optimization
  • 4.1 Direct Optimization
  • 4.2 Greedy Optimization for Extractive Summarization
  • 4.3 Reranking
  • 5 Analysis
  • 5.1 Approximate Inference Effectiveness
  • 5.2 Undesirable Metric Biases
  • 5.3 Reference-Free Metrics as Pseudo-References
  • 6 Discussion
  • 6.1 Reference-Free Evaluation
  • 6.2 What about Inherently Reference-Free Evaluations?
  • 7 Related Work
  • 8 Conclusion
  • Acknowledgments
  • Limitations
  • References
  • A Implementation Details
  • B Additional Results

Knowls

  1. Knowl 1 — Reference-free metrics are conditional generation models

    theoretical result

    A conditional text-generation model can be represented by a scoring function θ(x,y)∈R\theta(x,y)\in\mathbb{R}, where x∈Xx\in\mathcal{X} is an input text and y∈Yy\in\mathcal{Y} is a candidate output text. Its inference procedure selects a high-scoring output according to

    fθ(x)=arg⁡max⁡y∈Yθ(x,y).f_\theta(x)=\arg\max_{y\in\mathcal{Y}}\theta(x,y).

    A reference-based metric instead scores an output using an additional human-written reference y∗y^*, as Mref(x,y,y∗)M_{\mathrm{ref}}(x,y,y^*), whereas a reference-free metric scores only the input and candidate, as Mfree(x,y)M_{\mathrm{free}}(x,y). Because a reference-free metric has the same input dependence as a conditional generation model, it can itself be treated as such a model, whether or not its developers explicitly describe it that way. Consequently, an inference procedure always exists in principle that finds the metric-optimal output:

    gMfree(x)=arg⁡max⁡y∈YMfree(x,y).g_{M_{\mathrm{free}}}(x)=\arg\max_{y\in\mathcal{Y}}M_{\mathrm{free}}(x,y).

    This procedure may be computationally impractical, but the output maximizing the reference-free score is already defined by the metric. Thus, the metric's underlying model is the theoretically best-performing model according to that metric.

  2. Knowl 2 — Test-time optimization procedures for reference-free metrics

    algorithm

    The paper operationalizes reference-free metrics as test-time objectives using three approximate inference procedures. Each procedure takes an input source text or document and a reference-free metric and returns an output selected to maximize that metric, without using a human-written reference.

    • Direct optimization for Prism-src: Prism-src scores a translation by its average log-probability under a learned sequence-to-sequence machine-translation model conditioned on the source. Beam search with that same translation model therefore produces an approximate Prism-src-optimal translation.

    • Greedy extractive optimization for QuestEval: For a document containing sentences s1,…,sns_1,\ldots,s_n and a target summary length of LL sentences, start with an empty summary. At each step, score the summary formed by adding each not-yet-selected sentence, add the sentence producing the largest QuestEval increase, and stop after LL sentences. Exhaustively enumerating all LL-sentence summaries would find the exact optimum but is generally expensive; the greedy procedure approximates it directly at test time.

    • Reranking for any of the analyzed metrics: Run standard beam search with a task-specific pretrained sequence-to-sequence model using beam size KK. At every decoding step, retain the KK partial outputs with the highest pretrained-model log-likelihoods. At completion, score the resulting KK complete outputs with the reference-free metric and return the highest-scoring candidate. For summarization, the task model was BART trained on CNN/DailyMail; for machine translation, it was Facebook's WMT'19 submission for English–German, German–English, English–Russian, and Russian–English.

    The procedures restrict the search space differently, but all make it possible to optimize a reference-free evaluation score after training and during inference.

  3. Knowl 3 — Experimental comparison of three reference-free metrics

    experimental setup

    The analysis evaluates Prism-src and COMET-QE for machine translation and QuestEval for summarization. Prism-src assigns a translation a conditional log-probability under a multilingual sequence-to-sequence translation model. COMET-QE uses the COMET cross-lingual encoder architecture but forms its representation from only the source and candidate translation, omitting the reference. QuestEval generates question–answer pairs from the source document and generated summary, then scores the summary by how many questions can be answered correctly in the opposite text; the experiments use its learned query-weighting component.

    Machine-translation experiments use the WMT'19 metrics shared-task data, containing human-written reference translations and human-judged outputs from 10–20 translation systems across 18 language pairs. Summarization experiments use SummEval, with outputs from 16 models, and REALSumm, with outputs from 25 models; both datasets are based on CNN/DailyMail and include reference summaries and human-judged system outputs.

    The authors compare the reference-free scores with reference-based metrics used as quality indicators: BLEURT, BLEU, and BERTScore for translation, and ROUGE, BERTScore, and QAEval for summarization. The reference-based scores are treated as proxies for human quality when the newly generated outputs are compared with the existing systems, because the newly generated outputs were not newly judged by humans.

  4. Knowl 4 — Simple inference procedures produce the highest reference-free scores but not high-quality text

    empirical result

    Across every machine-translation language pair and both summarization datasets, the proposed approximate inference procedures produced outputs with the highest scores under the corresponding reference-free metric among the systems being compared. For example, reranking WMT'19 German-to-English translations by COMET-QE increased the best system score from 0.3470.347 to 0.4780.478, a relative improvement of 38%.

    The gains in reference-free score did not consistently indicate better text under reference-based quality indicators. Prism-src direct optimization produced the highest Prism-src score but only an average BLEURT score among translation systems. Greedy extractive optimization of QuestEval on REALSumm produced summaries among the lowest-performing outputs according to ROUGE-2. Thus, a reference-free metric can be optimized effectively even when the resulting text is not correspondingly strong under reference-based measures.

  5. Knowl 5 — Reference-free metrics favor learned-model outputs over human references

    empirical result

    The experiments show that reference-free metrics can assign higher scores to model-generated outputs than to human-written reference texts, even when the human text is expected to be higher quality. In every analyzed setting, outputs deliberately optimized for the reference-free metric scored above the corresponding human references. Other systems that had not directly optimized the metric also frequently scored above the references, especially for Prism-src and QuestEval; this effect was less pronounced for COMET-QE. For German-to-English translation, only one WMT'19 system had a lower Prism-src score than the human reference.

    The reported translation examples illustrate that this is not merely a harmless preference. For a German source where Kater means hangover in context, the human reference uses “hangover” and receives a Prism-src score of −1.6-1.6, while the incorrect candidate “one powerful cat” receives −0.4-0.4. For another source containing the idiom mit Mann und Maus, the human reference correctly conveys “with all means available” and receives −4.8-4.8, while the literal and incorrect “with man and mouse” receives −0.4-0.4. The metric therefore rewards outputs resembling its learned translation behavior rather than necessarily rewarding human-quality meaning.

  6. Knowl 6 — Reference-free evaluation is strongly correlated with evaluation against a model-generated pseudo-reference

    empirical result

    For each reference-free metric, the paper defines a pseudo-reference as an output produced by optimizing that metric: a Prism-src translation, a COMET-QE reranked translation, or a QuestEval-optimized summary. Other systems are then evaluated against this pseudo-reference using a reference-based metric. At the system level, these pseudo-reference scores are strongly correlated with the original reference-free scores.

    Using BERTScore against Prism-src pseudo-references yields an average Pearson correlation of 0.950.95 with Prism-src scores across the translation settings. The corresponding average correlation for BERTScore against COMET-QE pseudo-references is 0.920.92. For summarization, QAEval against QuestEval pseudo-references has an average Pearson correlation of 0.880.88 across SummEval and REALSumm. Many individual correlations are at least 0.900.90.

    These results indicate that a reference-free metric behaves roughly like a reference-based metric whose gold reference is the metric's own best or near-best output. The metric consequently favors its own output and other outputs that are similar to it, rather than independently measuring closeness to human-quality text.

  7. Knowl 7 — Pseudo-reference evaluation penalizes outputs that surpass the metric's own model

    limitation

    Viewing a reference-free metric as a pseudo-reference evaluator exposes an intrinsic ceiling on what it can recognize. If a system produces a translation or summary that is genuinely higher quality than the metric's optimized pseudo-reference, differences from that pseudo-reference can lower the system's score even when those differences improve the text. The reference-free metric therefore gives misleading evaluations whenever a system is better than the metric's underlying model in ways that the metric does not represent.

    This limitation follows from treating the metric's own output as the gold standard: similarity to that output is rewarded, while improvements that move away from it can be penalized. The metric is consequently unsuitable for reliably ranking systems whose quality may exceed the quality or stylistic behavior of the metric's underlying model.

  8. Knowl 8 — The suspected source of the learned-model bias

    model/method

    The authors hypothesize that the preference for learned-model outputs arises from how the internal components of the metrics are trained. Prism-src uses a machine-translation model trained on standard machine-translation data. QuestEval's query-weighting model predicts whether questions are answered in CNN/DailyMail reference summaries and is trained using the same dataset family on which the evaluated summarization systems were trained.

    Models trained on the same task data can learn similar signals, behaviors, and mistakes, so the internal models of Prism-src and QuestEval may identify as high quality the same patterns that task-specific generation systems learn to produce. Human-written text may not exhibit those model-specific signals and can therefore receive lower scores. The authors expect this bias to be less severe for COMET-QE because COMET-QE is trained on manually collected human quality judgments, including judgments of human-written translations, rather than being derived directly from a generation model.

  9. Knowl 9 — Reference-free metrics should diagnose model behavior, not measure progress

    limitation

    The paper recommends against using reference-free metric scores to declare that one machine-translation or summarization system is better than another. Since the metric's underlying model is already optimal under the metric and simple inference procedures can find high-scoring outputs, optimizing the score encourages systems to imitate the metric's model rather than improve toward human-quality text. Under this evaluation regime, progress can become progress in test-time metric optimization, and the attainable text quality is bounded by the quality and biases of the metric itself.

    When references are unavailable, the authors recommend investing in collecting human-written references where possible. Reference-free metrics can still be reported as diagnostic statistics—for example, a language-model perplexity can describe how likely a summary is under that language model—but the statistic should be interpreted only as measuring the behavior encoded by its underlying model. It should not be used as the primary objective for improving generation quality.

  10. Knowl 10 — Scope and evidential limitations of the analysis

    limitation

    The theoretical argument is intended to apply broadly to reference-free metrics, but the empirical evidence covers only one machine-translation dataset and two summarization datasets. The extent to which the observed optimization effects and biases appear in other generation tasks is therefore not established. The experiments also use reference-based metrics as surrogates for human judgments when assessing outputs created by the optimization procedures; direct human evaluation of those outputs was not performed.

    The paper does not provide a general alternative that automatically evaluates inherently reference-free properties such as fluency or summary faithfulness without introducing analogous model-dependent biases. It therefore establishes limitations of using such metrics as progress measures without resolving how all reference-free properties should instead be evaluated.

Coverage note — Detailed appendix plots and repeated combinations of optimization and reference-based metrics were omitted because they reproduce the same optimization, bias, and pseudo-reference findings rather than adding distinct contributed claims.

References

  1. 1.Sweta Agrawal, George Foster, Markus Freitag, and Colin Cherry. 2021. Assessing Reference-Free Peer Evaluation for Machine Translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1158–1171, Online. Association for Computational Linguistics.
  2. 2.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating Evaluation in Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online. Association for Computational Linguistics.
  3. 3.Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021a. Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary. Transactions of the Association for Computational Linguistics, 9:774–789.
  4. 4.Daniel Deutsch, Rotem Dror, and Dan Roth. 2021b. A Statistical Analysis of Summarization Evaluation Metrics Using Resampling Methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
  5. 5.Daniel Deutsch and Dan Roth. 2020. SacreROUGE: An Open-Source Library for Using and Developing Summarization Evaluation Metrics. In Proceedings of Second Workshop for NLP Open Source Software (NLP-OSS), pages 120–125, Online. Association for Computational Linguistics.
  6. 6.Daniel Deutsch and Dan Roth. 2022. Repro: An Open-Source Library for Improving the Reproducibility and Usability of Publicly Available Research Code. ArXiv, abs/2204.13848.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Esin Durmus, Faisal Ladhak, and Tatsunori Hashimoto. 2022. Spurious Correlations in Reference-Free Evaluation of Text Generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1443–1454, Dublin, Ireland. Association for Computational Linguistics.
  9. 9.Alexander Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, and Dragomir Radev. 2021. SummEval: Re-evaluating Summarization Evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  10. 10.Fern, Patrick es, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022. Quality-Aware Decoding for Neural Machine Translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1396–1412, Seattle, United States. Association for Computational Linguistics.
  11. 11.Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019. Findings of the WMT 2019 Shared Tasks on Quality Estimation. In Proceedings of the Fourth Conference on Machine Translation (Volume 3: Shared Task Papers, Day 2), pages 1–12, Florence, Italy. Association for Computational Linguistics.
  12. 12.Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022. High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics. Transactions of the Association for Computational Linguistics, 10:811–825.
  13. 13.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021. Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  14. 14.Yang Gao, Wei Zhao, and Steffen Eger. 2020. SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1347–1354, Online. Association for Computational Linguistics.
  15. 15.Yvette Graham. 2015. Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 128–137, Lisbon, Portugal. Association for Computational Linguistics.
  16. 16.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7514–7528, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7856–7870, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Reno Kriz, Marianna Apidianaki, and Chris Callison-Burch. 2020. Simple-QE: Better Automatic Quality Estimation for Text Simplification. arXiv preprint arXiv:2012.12382.
  19. 19.Shankar Kumar and William Byrne. 2004. Minimum Bayes-Risk Decoding for Statistical Machine Translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 169–176, Boston, Massachusetts, USA. Association for Computational Linguistics.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pretraining for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  21. 21.Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  22. 22.Hui Lin and Jeff Bilmes. 2011. A Class of Submodular Functions for Document Summarization. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 510–520, Portland, Oregon, USA. Association for Computational Linguistics.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mar, Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, abs/1907.11692.
  24. 24.Annie Louis and Ani Nenkova. 2013. Automatically Assessing Machine Summary Content Without a Gold Standard. Computational Linguistics, 39(2):267–300.
  25. 25.Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019. Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 62–90, Florence, Italy. Association for Computational Linguistics.
  26. 26.Louis Martin, Samuel Humeau, Pierre-Emmanuel Mazaré, Éric de La Clergerie, Antoine Bordes, and Benoît Sagot. 2018. Reference-less Quality Estimation of Text Simplification Systems. In Proceedings of the 1st Workshop on Automatic Text Adaptation (ATA), pages 29–38, Tilburg, the Netherlands. Association for Computational Linguistics.
  27. 27.Shikib Mehri and Maxine Eskenazi. 2020. USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 681–707, Online. Association for Computational Linguistics.
  28. 28.Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017. SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents. In Proc. of the Conference on Artificial Intelligence (AAAI).
  29. 29.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çaglar Gulçehre, and Bing Xiang. 2016. Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  30. 30.Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019. Facebook FAIR’s WMT19 News Translation Task Submission. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 314–319, Florence, Italy. Association for Computational Linguistics.
  31. 31.Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alex Fraser, Shankar Kumar, Libin Shen, David Smith, Katherine Eng, Viren Jain, Zhen Jin, and Dragomir Radev. 2004. A Smorgasbord of Features for Statistical Machine Translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 161–168, Boston, Massachusetts, USA. Association for Computational Linguistics.
  32. 32.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  33. 33.Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels. Association for Computational Linguistics.
  34. 34.Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André FT Martins, and Alon Lavie. 2021. Are References Really Needed? Unbabel-IST 2021 Submission for the Metrics Shared Task. In Proceedings of the Sixth Conference on Machine Translation, Online. Association for Computational Linguistics.
  35. 35.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A Neural Framework for MT Evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  36. 36.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization Asks for Fact-based Evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019. Answers Unite! Unsupervised Metrics for Reinforced Summarization Models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3246–3256, Hong Kong, China. Association for Computational Linguistics.
  38. 38.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  39. 39.Libin Shen, Anoop Sarkar, and Franz Josef Och. 2004. Discriminative Reranking for Machine Translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 177–184, Boston, Massachusetts, USA. Association for Computational Linguistics.
  40. 40.Lucia Specia, Frédéric Blain, Marina Fomicheva, Erick Fonseca, Vishrav Chaudhary, Francisco Guzmán, and André F. T. Martins. 2020. Findings of the WMT 2020 Shared Task on Quality Estimation. In Proceedings of the Fifth Conference on Machine Translation, pages 743–764, Online. Association for Computational Linguistics.
  41. 41.Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021. Findings of the WMT 2021 Shared Task on Quality Estimation. In Proceedings of the Sixth Conference on Machine Translation, pages 684–725, Online. Association for Computational Linguistics.
  42. 42.Brian Thompson and Matt Post. 2020. Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 90–121, Online. Association for Computational Linguistics.
  43. 43.Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020. Fill in the BLANC: Human-free quality estimation of document summaries. In Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems, pages 11–20, Online. Association for Computational Linguistics.
  44. 44.Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, and Ion Androutsopoulos. 2019. SUM-QE: a BERT-based Summary Quality Estimation Model. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6005–6011, Hong Kong, China. Association for Computational Linguistics.
  45. 45.Shiyue Zhang and Mohit Bansal. 2021. Finding a Balanced Degree of Automation for Summary Evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6617–6632, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.

Citation

MLA
Deutsch, D., et al. “On the Limitations of Reference-Free Evaluations of Generated Text”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 10960–77, https://doi.org/10.18653/v1/2022.emnlp-main.753.
APA
Deutsch, D., Dror, R., & Roth, D. (2022). On the Limitations of Reference-Free Evaluations of Generated Text. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10960–10977. https://doi.org/10.18653/v1/2022.emnlp-main.753
Chicago
Deutsch, D., R. Dror, and D. Roth. 2022. “On the Limitations of Reference-Free Evaluations of Generated Text”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10960–77. https://doi.org/10.18653/v1/2022.emnlp-main.753.
Harvard
Deutsch, D., Dror, R. and Roth, D. (2022) “On the Limitations of Reference-Free Evaluations of Generated Text”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10960–10977. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.753.
Vancouver
1. Deutsch D, Dror R, Roth D (2022) On the Limitations of Reference-Free Evaluations of Generated Text. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10960–10977

BibTeX

@inproceedings{deutsch-etal-2022-limitations,
    title = "On the Limitations of Reference-Free Evaluations of Generated Text",
    author = "Deutsch, Daniel  and
      Dror, Rotem  and
      Roth, Dan",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.753/",
    doi = "10.18653/v1/2022.emnlp-main.753",
    pages = "10960--10977"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/