On the Evaluation Metrics for Paraphrase Generation
Lingfeng ShenLemao LiuHaiyun JiangShuming Shi
Proposes ParaScore, a paraphrase evaluation metric that explicitly incorporates lexical divergence to significantly improve correlation with human judgments over existing automatic metrics.
Paraphrase generation is a core capability in language technology, supporting applications such as writing assistants, question answering, and machine translation. Evaluating whether a machine-generated paraphrase is effective requires measuring two distinct qualities: semantic similarity, which ensures the original meaning is preserved, and lexical divergence, which ensures the wording is varied rather than copied. Despite rapid algorithmic progress, standard evaluation metrics borrowed from translation and summarization tasks fail to reliably measure these criteria, creating a significant evaluation gap.
The article evaluates the reliability of common automatic metrics against human judgment, investigates why standard evaluation practices fall short, and introduces ParaScore, a new evaluation metric designed to measure both meaning preservation and wording variation accurately.
To investigate metric performance, the authors conducted statistical correlation analyses on English and Chinese datasets consisting of 761 and 550 input sentences with over 12,000 candidate paraphrases. They evaluated established n-gram and embedding-based metrics in both reference-based formats (comparing output to a human reference) and reference-free formats (comparing output directly to the input). They also applied attribution analysis to isolate the specific contributions of semantic similarity and lexical divergence to overall human quality scores.
The investigation produced four central findings. First, widely used metrics align poorly with human judgments; for example, standard word-overlap metrics like BLEU showed near-zero or even negative correlations with human annotations. Second, reference-free metrics consistently outperformed reference-based versions because typical test candidates share closer lexical proximity to the input text than to a single human reference. Third, existing metrics capture semantic similarity reasonably well but fail entirely to reward lexical divergence, frequently assigning high scores to verbatim copies of the input. Fourth, human evaluation rewards lexical variation only up to a point; beyond a moderate threshold (an edit distance around 0.35), additional wording variation provides no further quality benefit. Incorporating these insights, the proposed metric, ParaScore, achieved the highest human alignment, improving correlation scores over standard embedding metrics across standard and stress-tested benchmarks.
These findings indicate that relying on legacy metrics like BLEU or standard ROUGE creates significant risk in product development and benchmarking by misjudging paraphrase quality and penalizing creative, valid phrasing. Furthermore, the analysis reveals that standard evaluation benchmarks themselves are skewed, as they inadequately penalize copied text and lack natural diversity. Adopting an evaluation approach that incorporates a bounded threshold for wording variation provides a more accurate assessment of natural language generation quality.
Organizations developing or deploying paraphrasing systems should transition away from legacy translation metrics toward evaluation frameworks that explicitly balance semantic preservation with threshold-based lexical divergence, such as ParaScore. Development teams must also build more representative benchmark datasets that include natural lexical variation to avoid overfitting models to superficial word matching.
A current limitation of this work is that the stress-test benchmarks introduced variation by artificially copying input sentences rather than sourcing natural, diverse human phrasing. Nevertheless, confidence in the primary findings remains high given the consistent mathematical and empirical results across multiple metric families and two distinct languages.
- Paper: Bleu: a Method for Automatic Evaluation of Machine Translation, Kishore Papineni et al. (2002). Read BLEU’s original reference-based n-gram design first to understand the legacy metric whose paraphrase-evaluation failures this paper diagnoses.
- Paper: Collecting Highly Parallel Data for Paraphrase Evaluation, David L. Chen et al. (2011). Its PINC metric operationalizes lexical divergence alongside BLEU for paraphrases, providing a direct precursor to this paper’s effort to balance variation with meaning preservation.
- Paper: On the Blind Spots of Model-Based Evaluation Metrics for Text Generation, Tianxing He et al. (2023). Building on the problem of metrics misjudging generated text, this later work stress-tests model-based evaluators for blind spots beyond paraphrase meaning and lexical variation.
- Paper: LENS: A Learnable Evaluation Metric for Text Simplification, Mounica Maddela et al. (2023). This later metric study carries the copying-versus-valid-variation problem into text simplification, developing a learned evaluator that accounts for task-specific changes.
