On the Blind Spots of Model-Based Evaluation Metrics for Text Generation

Tianxing HeJingyu ZhangTianle WangSachin KumarKyunghyun ChoJames R. GlassYulia Tsvetkov

article2023ACL72 citations

Exposes critical failure modes in popular pretrained language model-based evaluation metrics like BERTScore and MAUVE using synthetic stress tests, while providing practical workarounds to ensure more reliable text generation assessment.

Listen

Automated evaluation metrics powered by pretrained language models have gained widespread adoption across natural language generation tasks such as machine translation, text summarization, and open-ended generation. While these model-based metrics demonstrate strong statistical correlation with human judgments on standard benchmarks, the article addresses an urgent and under-explored risk: underlying model limitations and architectural design choices introduce critical blind spots that allow severely degraded or manipulated text to receive deceptively high scores.

The main objective of the article is to systematically analyze the robustness of widely used evaluation metrics—including BERTScore, BARTScore, MAUVE, COMET, and UniEval—by developing a stress-testing framework that introduces synthetic linguistic, structural, and adversarial errors into human-written text.

The researchers conducted controlled experiments across established benchmarks: WikiText-103 for open-ended generation, CNN/DailyMail for multi-reference summarization, and WMT21 alongside a newly curated paragraph-level TED-Talks dataset for translation. The evaluation protocol established a baseline score on clean, human-authored text and introduced controlled perturbations across 18 error types covering fluency, consistency, positional bias, and adversarial injection. If a metric failed to assign a lower score to the degraded text than to the clean gold standard, or failed to decrease monotonically as the noise level increased, it was deemed to have failed the test.

The evaluation revealed several critical failures across leading metrics. First, several summarization metrics failed truncation tests; BERTScore's f-measure and BARTScore variants did not penalize severely truncated texts because precision scores increased as summaries were cut short, offsetting drops in recall. Second, MAUVE configured with its default GPT-2 representations exhibited extreme positional insensitivity, showing only a 1.3% to 6.5% score drop when errors were placed at the beginning or middle of paragraphs, compared to over a 95% drop when identical errors occurred at the end. Third, question-answering evaluators like UniEval were easily manipulated by adversarial text injection, awarding higher overall scores (0.905 versus the gold 0.864) to meaningless phrases that explicitly asserted the summary was high quality. Fourth, probability-based evaluators exhibited strong self-evaluation bias: systems evaluated by their own underlying model architecture (such as BART evaluating BART, or GPT evaluating GPT) received unfairly favorable scores over superior or larger alternative architectures. Finally, multiple metrics preferred repetitive phrases, and several reference-free or recall-oriented metrics awarded higher scores to direct copies of the full source text than to actual human summaries.

These findings have direct operational and governance implications for deploying natural language generation systems. Relying on single, unvetted automatic metrics creates substantial risks of selecting degraded models, misinterpreting production performance, or exposing competitive leaderboards to gaming and prompt injection. Because flawed metrics can mask severe content loss, hallucinations, and temporal incoherence, decision-makers cannot depend exclusively on standard automated leaderboards for critical deployments.

The article recommends several immediate countermeasures for practitioners and developers. Metric users should avoid single-metric evaluations and instead report complementary metrics, pairing reference-based and probability-based tools with diversity measures such as n-gram repetition penalties. For summarization, precision, recall, and f-measure should all be reported rather than relying solely on f-measure. Generation systems must not be evaluated using the exact same underlying model family to avoid self-evaluation bias. For open-ended generation, switching MAUVE representations from GPT-2 to ELECTRA-large or RoBERTa-large substantially restores sensitivity to positional and sentence-order errors. Finally, contest organizers should deploy explicit filtering to detect source copying and length anomalies.

These conclusions are supported by controlled, reproducible experiments across multiple random seeds, but certain limitations remain. The diagnostic stress tests rely primarily on synthetic perturbations rather than real-world machine error distributions, and the empirical evaluations were conducted exclusively on English datasets. Further validation is required for low-resource and multilingual environments, as well as specialized tasks like dialogue and factuality checking.

Cover for On the Blind Spots of Model-Based Evaluation Metrics for Text Generation

Abstract

In this work, we study the blind spots of model-based evaluation metrics for text generation. We first analyze the behaviors of model-based metrics under adversarial perturbations and find that they are vulnerable to adversarial attacks. We then show that the blind spots are caused by the fact that model-based metrics are trained on the same data distribution as the generation models. We further propose a simple method to mitigate the blind spots by training the metric on a different data distribution. Experiments on multiple text generation tasks demonstrate the effectiveness of our method.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 3 Tasks and Datasets
  • 4 Metrics
  • 5 Stress Tests and Results
  • 5.1 The Positioned Error Test
  • 5.2 The Injection Test
  • 5.3 The Frequent n-gram Test
  • 5.4 The Self-Evaluation Bias
  • 5.5 Fluency & Consistency Tests
  • 5.5.1 Noise Types and Setup
  • 5.5.2 Results
  • 6 Discussion
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • Supplemental Materials
  • A Implementation Details of Metrics or Tests
  • B More Information on Datasets
  • B.1 The TED-MT Dataset
  • B.2 WikiText Preprocessing
  • C Details on the Positioned Error Test
  • C.1 Auxiliary Results
  • C.2 Attention Pattern Analysis
  • C.3 MAUVE Correlation with Human Judgment
  • D The Copy-Source Test
  • E The Repetition Test
  • F Auxiliary Results for the Injection Test
  • G Auxiliary Results for the Frequent n-gram Test
  • H Details on the Finetuning (Self-Evaluation)
  • I Auxiliary Description and Results of the Fluency and Consistency Tests
  • J Can We Automate the Detection?
  • J.1 Attack Algorithm Details

Knowls

  1. Knowl 1 — Synthetic stress-test protocol and evaluation scope

    model/method

    The paper evaluates text-generation metrics by starting with relatively clean human-written outputs, applying controlled synthetic errors to create noised hypotheses, and checking whether metric scores respond in the expected direction. In a multi-reference translation setting, one human translation is treated as the gold hypothesis and another as the reference; the source and reference remain unchanged when noise is applied. A metric fails a test when the noised hypotheses are not scored below the gold hypotheses. When a noise type has levels, scores are expected to decrease as the noise increases.

    For fluency and consistency tests, the authors use the noise-ratio as a rough measure of perturbation size: 1∣H∣∑h∈HLev⁡(h′,h)len⁡(h)\frac{1}{|H|}\sum_{h\in H}\frac{\operatorname{Lev}(h',h)}{\operatorname{len}(h)}, where HH is the set of gold hypotheses, hh is one gold hypothesis, h′h' is its perturbed version, Lev⁡\operatorname{Lev} is token-level Levenshtein distance, and len⁡(h)\operatorname{len}(h) is the number of tokens in hh. For sentence-switching noise, they divide this ratio by two because Levenshtein distance does not represent a sentence switch as a single operation. The broader suite contains 10 fluency tests and 8 consistency tests, including truncation, token and function-word edits, sentence switching or replacement, negation, and entity or verb substitutions.

    The experiments cover open-ended generation, summarization, and translation. Open-ended generation uses 2,000 WikiText-103 paragraphs of about 256 tokens, split into 1,000 gold hypotheses and 1,000 references. Summarization uses 100 CNN/DailyMail test examples, with the dataset summaries as gold hypotheses and 10 additional human summaries per example as references. Translation uses 1,000 German–English WMT21 pairs with one human translation as the gold hypothesis and the other as reference; the authors also build TED-MT, a set of 100 Chinese–English paragraphs averaging about seven sentences, to test paragraph-level errors. Metrics include PLM-based similarity, likelihood, and evaluator metrics, as well as BLEU or ROUGE baselines. The diagnostics are synthetic and English-only, and passing them is not evidence that a metric has no other blind spots.

  2. Knowl 2 — MAUVE-GPT2 misses errors near the beginning and middle of text

    empirical result

    The positioned-error test replaces a contiguous span of 10 tokens in each WikiText gold hypothesis with either random vocabulary tokens or a random permutation of the original span. The span is placed at the beginning, middle, or end. Because MAUVE represents texts using a pretrained model's final-token representation, this test probes whether errors at different positions affect its distributional score.

    MAUVE-GPT2 has a gold score of 0.961, but barely penalizes several beginning- and middle-position errors: random replacement scores 0.949 at the beginning (a 1.3% decrease) and 0.898 in the middle (6.5%); shuffling scores 0.916 at the beginning (4.7%) and 0.943 in the middle (1.8%). It does sharply penalize end-position errors: random replacement scores 0.005 (99.4% decrease), and shuffling scores 0.020 (97.9%). MAUVE-RoBERTa, whose gold score is 0.969, responds much more strongly to the same beginning and middle errors: random replacement scores 0.037 at the beginning and 0.100 in the middle, while shuffling scores 0.342 and 0.603, respectively. The authors associate GPT-2's positional blind spot with its attention concentrating on nearby context; their attention analysis is evidence for this explanation, not a proof of cause. They recommend considering RoBERTa or ELECTRA features rather than relying on MAUVE-GPT2 alone.

  3. Knowl 3 — Language-model perplexity rewards random sequences of frequent n-grams

    empirical result

    For open-ended generation, the authors create 256-token hypotheses by uniformly concatenating frequent four-grams collected from WikiText. These sequences are intended to be locally familiar but globally incoherent. Scores are negated perplexities, so a higher (less negative) score indicates that the metric prefers the text.

    GPT-PPL gives the gold text a score of -25.640, but scores sequences made from the top 10, top 50, and top 100 frequent four-grams at -4.456, -11.640, and -18.160, respectively. MLM-PPL likewise scores the gold at -2.994 and the three synthetic sets at -1.139, -2.469, and -3.971. Thus both likelihood-based metrics can prefer nonsensical concatenations of frequent phrases to the gold text, especially when the phrases are drawn from the most frequent n-grams. The authors observe that high next-token probabilities often cluster at the end of each n-gram, consistent with a reliance on local context. They do not observe the same problem in their translation or summarization tests, which they suggest may be because random n-grams align poorly with source or reference texts. The result motivates pairing likelihood scores with diversity measures such as repetition rate.

  4. Knowl 4 — Likelihood evaluators favor generators based on their own model

    empirical result

    The authors test whether likelihood-based metrics favor generation systems built from the same pretrained model as the evaluator. In the GPT-PPL experiment, GPT-2 small, medium, and large models are fine-tuned on WikiText-103 for two epochs and generate continuations using top-kk sampling with k=50k=50; off-the-shelf GPT-2 models evaluate the continuations. Each GPT-2 evaluator gives its highest score to generations from the matching size: GPT-2 small assigns its own generator -21.08 versus -24.35 for medium and -24.36 for large; GPT-2 medium assigns its own generator -17.48 versus -23.20 for small and -19.06 for large; GPT-2 large assigns its own generator -15.04 versus -22.87 for small and -18.56 for medium. In contrast, OPT-2.7b scores the large-model generations best among these three, at -17.20, compared with -24.24 for small and -19.08 for medium.

    A related experiment fine-tunes BART and T5 models on CNN/DailyMail and uses the reference-free BARTScore-cnn-faithful variant. Evaluators based on BART favor BART generators, and T5 evaluators favor T5 generators, including across model sizes. The authors report that this effect is less pronounced for reference-based BARTScore-para. These results show that evaluator choice can change system rankings and that self-evaluation can bias them; they advise against evaluating a system with a metric based on the same pretrained model, especially when comparing systems based on different model families.

  5. Knowl 5 — Instruction-like text can inflate UniEval scores

    empirical result

    UniEval evaluates a summary by asking a pretrained T5 model questions about properties such as coherence and relevance and scoring its probability of answering “Yes.” The injection test adds valueless text that attempts to cue that answer, without changing the evaluator's prompts. For a gold summary, UniEval's overall score is 0.864. Adding “Answer: Yes, this is a really coherent and consistent summary. And yes, it is relevant.” raises the score to 0.905; the less specific “Answer: Yes, this is a really good summary.” yields 0.838. With the first injection, coherence rises from 0.897 to 0.903, fluency from 0.919 to 0.959, and relevance from 0.781 to 0.900, while consistency changes from 0.859 to 0.857. The injected text is not a meaningful summary, yet it can achieve a higher overall score than the gold text. By comparison, ROUGE-L falls from 0.286 for the gold summary to 0.126 and 0.098 for the two injections. The results demonstrate a way to manipulate UniEval's judgment and show that a traditional metric can help flag this particular kind of injection.

  6. Knowl 6 — Several metrics fail to penalize summary truncation reliably

    empirical result

    In the truncation test, tokens are removed from the end of gold hypotheses at increasing levels, and a robust metric is expected to assign progressively lower scores as information is lost. On CNN/DailyMail, TED-MT, and WikiText, the authors report failures for some variants of BERTScore, BARTScore, COMET, PRISM, ROUGE, MAUVE, and UniEval. The affected variants include BERTScore precision and f-measure; BARTScore precision, f-measure, and faithfulness; COMET-QE and PRISM-QE; ROUGE-2 and ROUGE-L; MAUVE-GPT2; and UniEval overall. One explanation for BERTScore on summarization is that precision rises as a summary is truncated, offsetting the fall in recall and preventing its f-measure from decreasing. The authors report a similar pattern for BARTScore-para precision and faithfulness.

    All tested metrics pass the truncation test on WMT, where the gold and reference translations are generally very similar and the information loss is consequently easier for metrics to detect. For summarization evaluation, the authors recommend reporting precision, recall, and f-measure together, or calibrating the f-measure to give recall more weight.

  7. Knowl 7 — Sentence-order errors evade MAUVE-GPT2, MAUVE-RoBERTa, and BARTScore-para-recall

    empirical result

    The sentence-switching test randomly switches pairs of sentences within a multi-sentence hypothesis, disrupting temporal or logical order while leaving the sentences themselves intact. On WikiText, MAUVE-GPT2 and MAUVE-RoBERTa do not consistently lower their scores as sentence switching increases. This is notable because WikiText hypotheses usually contain several sentences, so the perturbation can substantially disrupt paragraph-level organization. The test avoids switching the final sentence in these paragraphs to separate this result from MAUVE's sensitivity to errors at particular token positions. BARTScore-para-recall also fails the sentence-switching test on summarization.

    MAUVE-ELECTRA passes the reported sentence-switching test, as well as the other fluency and consistency tests in the suite. The authors suggest that ELECTRA's discriminative training may help it detect these errors. Because MAUVE-RoBERTa can miss temporal or logical disorder, they recommend combining it with another metric such as GPT-PPL rather than relying on it alone.

  8. Knowl 8 — Copying the source can outperform gold text on several metrics

    empirical result

    The copy-source test submits the source text itself as the hypothesis for translation or summarization, although the intended output is a translation or a summary. On WMT, COMET-QE scores the copied source at 0.126 versus 0.114 for the gold translation; on TED-MT the corresponding scores are 0.073 and 0.062. On summarization, BERTScore-recall scores the copied source at 0.332 versus 0.266 for the gold summary, and UniEval's overall score is 0.920 versus 0.864. Several BARTScore variants also score copied source text above the gold summary. These outcomes make source copying a potential evaluation loophole.

    The paper attributes the behaviors to different metric properties: COMET-QE does not check the hypothesis language ID; length-averaged BARTScore can fail to account for a hypothesis that is as long as the source article; recall-oriented BERTScore favors source coverage, an effect mitigated by using its f-measure; and UniEval's overall score can be misleadingly high even when its component judgments do not establish summary quality. Suggested safeguards include checking hypothesis/source similarity and expected output length for summarization, and checking language ID for translation.

  9. Knowl 9 — Repetition can improve scores from several likelihood and similarity metrics

    empirical result

    For the repetition test, the authors append 10, 20, or 30 copies of each gold hypothesis's final four-gram. Since repeated text degrades output quality, a robust metric should score these versions below the gold hypothesis. Instead, on WikiText, GPT-PPL rises from -21.81 for gold to -15.48, -10.70, and -8.080 for Rep-10, Rep-20, and Rep-30; MLM-PPL rises from -2.635 to -2.241, -2.019, and -1.867. Because these are negated perplexity scores, the increasingly higher values mean the metrics prefer the increasingly repetitive outputs. On summarization, BARTScore-cnn-precision similarly rises from -2.718 for gold to -2.122, -1.675, and -1.451 across these repetition levels. The paper finds the problem across GPT-PPL, MLM-PPL, and multiple BARTScore variants, although not every variant's scores change monotonically in every test. The authors recommend supplementing quality scores with diversity measures such as four-gram repetition or n-gram entropy.

  10. Knowl 10 — MAUVE-ELECTRA combines stronger stress-test behavior with higher human correlations

    empirical result

    The authors compare MAUVE variants using GPT-2, RoBERTa, and ELECTRA features. In their stress tests, MAUVE-ELECTRA is more sensitive than the other variants to positioned errors and passes the sentence-switching test that MAUVE-GPT2 and MAUVE-RoBERTa fail. They also reproduce a pairwise human-judgment evaluation on WebText and report Spearman correlations between MAUVE scores and fitted human preference rankings. For the human-like, interesting, and sensible judgments, respectively, correlations are 0.952, 0.738, and 0.881 for GPT-2; 0.929, 0.786, and 0.881 for RoBERTa; and 0.976, 0.857, and 0.976 for ELECTRA. ELECTRA therefore has the highest reported correlation on all three judgments as well as strong stress-test performance. The authors identify it as the best-performing MAUVE feature among those they tested, while noting that it penalizes some errors, such as article or preposition removal, especially sharply and may require calibration.

Coverage note — The detailed curves for the remaining individual fluency and consistency perturbations and the preliminary BERTScore adversarial-search pilot are omitted: the former are largely supporting cases for the general stress-test protocol, while the pilot is explicitly exploratory and its discovered transformations did not generalize across the WMT dataset.

References

  1. 1.Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021. Findings of the 2021 conference on machine translation (WMT21). In Proceedings of the Sixth Conference on Machine Translation, pages 1–88, Online. Association for Computational Linguistics.
  2. 2.Sriram Balasubramanian, Naman Jain, Gaurav Jindal, Abhijeet Awasthi, and Sunita Sarawagi. 2020. What’s in a name? are BERT named entity representations just as good for any other name? In Proceedings of the 5th Workshop on Representation Learning for NLP, pages 205–214, Online. Association for Computational Linguistics.
  3. 3.Yonatan Belinkov. 2022. Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics, 48(1):207–219.
  4. 4.Yonatan Belinkov and James Glass. 2019. Analysis Methods in Neural Language Processing: A Survey. Transactions of the Association for Computational Linguistics, 7:49–72.
  5. 5.Ozan Caglayan, Pranava Madhyastha, and Lucia Specia. 2020. Curious case of language generation evaluation metrics: A cautionary tale. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2322–2328, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  6. 6.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. ArXiv, abs/2006.14799.
  7. 7.Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021. Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2082–2095, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  8. 8.Minhao Cheng, Jinfeng Yi, Huan Zhang, Pin-Yu Chen, and Cho-Jui Hsieh. 2018. Seq2sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples. CoRR, abs/1803.01128.
  9. 9.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations.
  10. 10.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Yue Dong, Chandra Bhagavatula, Ximing Lu, Jena D. Hwang, Antoine Bosselut, Jackie Chi Kit Cheung, and Yejin Choi. 2021. On-the-fly attention modulation for neural generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1261–1274, Online. Association for Computational Linguistics.
  13. 13.Kevin Duh. 2018. The multitarget ted talks task. http://www.cs.jhu.edu/~kevinduh/a/multitarget-tedtalks/.
  14. 14.Allyson Ettinger. 2020. What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models. Transactions of the Association for Computational Linguistics, 8:34–48.
  15. 15.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  16. 16.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  17. 17.Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021. Robustness gym: Unifying the NLP evaluation landscape. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Demonstrations, pages 42–55, Online. Association for Computational Linguistics.
  18. 18.Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindřich Helcl, and Alexandra Birch. 2022. Survey of low-resource machine translation. Computational Linguistics, 48(3):673–732.
  19. 19.Mika Hämäläinen and Khalid Alnajjar. 2021. Human evaluation of creative NLG systems: An interdisciplinary survey on recent papers. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 84–95, Online. Association for Computational Linguistics.
  20. 20.Michael Hanna and Ondřej Bojar. 2021. A fine-grained analysis of BERTScore. In Proceedings of the Sixth Conference on Machine Translation, pages 507–517, Online. Association for Computational Linguistics.
  21. 21.Tianxing He and James Glass. 2019. Detecting egregious responses in neural sequence-to-sequence models. In International Conference on Learning Representations.
  22. 22.Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. 2021. Exposure bias versus self-recovery: Are distortions really incremental for autoregressive text generation? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5087–5102, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  23. 23.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. Advances in neural information processing systems, 28:1693–1701.
  24. 24.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7514–7528, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  26. 26.J. Edward Hu, Abhinav Singh, Nils Holzenberger, Matt Post, and Benjamin Van Durme. 2019. Large-scale, diverse, paraphrastic bitexts via sampling and clustering. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 44–54, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Jiabao Ji, Yoon Kim, James Glass, and Tianxing He. 2022. Controlling the focus of pretrained language generation models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3291–3306, Dublin, Ireland. Association for Computational Linguistics.
  28. 28.Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, and Noah A. Smith. 2022a. Transparent human evaluation for image captioning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3464–3478, Seattle, United States. Association for Computational Linguistics.
  29. 29.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, and Noah A. Smith. 2022b. Bidimensional leaderboards: Generate and evaluate language hand in hand. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3540–3557, Seattle, United States. Association for Computational Linguistics.
  30. 30.Marvin Kaster, Wei Zhao, and Steffen Eger. 2021. Global explainability of BERT-based evaluation metrics by disentangling along linguistic factors. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8912–8925, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Pei Ke, Hao Zhou, Yankai Lin, Peng Li, Jie Zhou, Xiaoyan Zhu, and Minlie Huang. 2022. CTRLEval: An unsupervised reference-free metric for evaluating controlled text generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2306–2319, Dublin, Ireland. Association for Computational Linguistics.
  32. 32.Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018. Sharp nearby, fuzzy far away: How neural language models use context. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 284–294, Melbourne, Australia. Association for Computational Linguistics.
  33. 33.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  34. 34.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  35. 35.Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. BERT-ATTACK: Adversarial attack against BERT using BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6193–6202, Online. Association for Computational Linguistics.
  36. 36.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  38. 38.John I Marden. 1995. Analyzing and modeling rank data. Chapman Hall, London.
  39. 39.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020. Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4984–4997, Online. Association for Computational Linguistics.
  40. 40.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  41. 41.Shikib Mehri and Maxine Eskenazi. 2020. USR: An unsupervised and reference free evaluation metric for dialog generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 681–707, Online. Association for Computational Linguistics.
  42. 42.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. CoRR, abs/1609.07843.
  43. 43.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  44. 44.Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018. Stress test evaluation for natural language inference. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2340–2353, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  45. 45.Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant. 2014. The CoNLL-2014 shared task on grammatical error correction. In Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task, pages 1–14, Baltimore, Maryland. Association for Computational Linguistics.
  46. 46.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  47. 47.Thang Pham, Trung Bui, Long Mai, and Anh Nguyen. 2021. Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? pages 1145–1160.
  48. 48.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. 2021. Mauve: Measuring the gap between neural text and human text using divergence frontiers. In Neural Information Processing Systems.
  49. 49.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  50. 50.Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. 2019. Perturbation sensitivity analysis to detect unintended model biases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5740–5745, Hong Kong, China. Association for Computational Linguistics.
  51. 51.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  52. 52.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  53. 53.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  54. 54.Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019. Are red roses red? evaluating consistency of question-answering models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6174–6184, Florence, Italy. Association for Computational Linguistics.
  55. 55.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  56. 56.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  57. 57.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Annual Meeting of the Association for Computational Linguistics.
  58. 58.Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. Bleurt: Learning robust metrics for text generation. In Proceedings of ACL.
  59. 59.Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi. 2022. On the evaluation metrics for paraphrase generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3178–3190, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  60. 60.Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. 2021. Findings of the WMT 2021 shared task on quality estimation. In Proceedings of the Sixth Conference on Machine Translation, pages 684–725, Online. Association for Computational Linguistics.
  61. 61.Ieva Staliunaitė and Ignacio Iacobacci. 2020. Compositional and lexical semantics in RoBERTa, BERT and DistilBERT: A case study on CoQA. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7046–7056, Online. Association for Computational Linguistics.
  62. 62.Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020. Assessing the benchmarking capacity of machine reading comprehension datasets. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8918–8927.
  63. 63.Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2022. Dialect-robust evaluation of generated text.
  64. 64.Brian Thompson and Matt Post. 2020. Automatic machine translation evaluation in many languages via zero-shot paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 90–121, Online. Association for Computational Linguistics.
  65. 65.Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63–76, Florence, Italy. Association for Computational Linguistics.
  66. 66.Doan Nam Long Vu, Nafise Sadat Moosavi, and Steffen Eger. 2022. Layer or representation space: What makes BERT-based evaluation metrics robust? In Proceedings of the 29th International Conference on Computational Linguistics, pages 3401–3411, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  67. 67.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  68. 68.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  69. 69.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  70. 70.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  71. 71.Wenda Xu, Yi-lin Tuan, Yujie Lu, Michael Saxon, Lei Li, and William Yang Wang. 2022. Not all errors are equal: Learning text generation metrics using stratified error synthesis. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  72. 72.Kevin Yang and Dan Klein. 2021. FUDGE: Controlled text generation with future discriminators. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3511–3535, Online. Association for Computational Linguistics.
  73. 73.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems, volume 34, pages 27263–27277. Curran Associates, Inc.
  74. 74.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pretrained transformer language models.
  75. 75.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  76. 76.Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018. Generating informative and diverse conversational responses via adversarial information maximization. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 1815–1825. Curran Associates, Inc.
  77. 77.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China. Association for Computational Linguistics.
  78. 78.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation. CoRR, abs/2210.07197.

Citation

MLA
He, T., et al. “On the Blind Spots of Model-Based Evaluation Metrics for Text Generation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12067–97, https://doi.org/10.18653/v1/2023.acl-long.674.
APA
He, T., Zhang, J., Wang, T., Kumar, S., Cho, K., Glass, J., & Tsvetkov, Y. (2023). On the Blind Spots of Model-Based Evaluation Metrics for Text Generation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12067–12097. https://doi.org/10.18653/v1/2023.acl-long.674
Chicago
He, T., J. Zhang, T. Wang, et al. 2023. “On the Blind Spots of Model-Based Evaluation Metrics for Text Generation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12067–97. https://doi.org/10.18653/v1/2023.acl-long.674.
Harvard
He, T. et al. (2023) “On the Blind Spots of Model-Based Evaluation Metrics for Text Generation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12067–12097. Available at: https://doi.org/10.18653/v1/2023.acl-long.674.
Vancouver
1. He T, Zhang J, Wang T, Kumar S, Cho K, Glass J, Tsvetkov Y (2023) On the Blind Spots of Model-Based Evaluation Metrics for Text Generation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12067–12097

BibTeX

@inproceedings{he-etal-2023-blind,
    title = "On the Blind Spots of Model-Based Evaluation Metrics for Text Generation",
    author = "He, Tianxing  and
      Zhang, Jingyu  and
      Wang, Tianle  and
      Kumar, Sachin  and
      Cho, Kyunghyun  and
      Glass, James  and
      Tsvetkov, Yulia",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.674/",
    doi = "10.18653/v1/2023.acl-long.674",
    pages = "12067--12097"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/