On the Evaluation Metrics for Paraphrase Generation

Lingfeng ShenLemao LiuHaiyun JiangShuming Shi

article2022EMNLP59 citations

Proposes ParaScore, a paraphrase evaluation metric that explicitly incorporates lexical divergence to significantly improve correlation with human judgments over existing automatic metrics.

Listen

Paraphrase generation is a core capability in language technology, supporting applications such as writing assistants, question answering, and machine translation. Evaluating whether a machine-generated paraphrase is effective requires measuring two distinct qualities: semantic similarity, which ensures the original meaning is preserved, and lexical divergence, which ensures the wording is varied rather than copied. Despite rapid algorithmic progress, standard evaluation metrics borrowed from translation and summarization tasks fail to reliably measure these criteria, creating a significant evaluation gap.

The article evaluates the reliability of common automatic metrics against human judgment, investigates why standard evaluation practices fall short, and introduces ParaScore, a new evaluation metric designed to measure both meaning preservation and wording variation accurately.

To investigate metric performance, the authors conducted statistical correlation analyses on English and Chinese datasets consisting of 761 and 550 input sentences with over 12,000 candidate paraphrases. They evaluated established n-gram and embedding-based metrics in both reference-based formats (comparing output to a human reference) and reference-free formats (comparing output directly to the input). They also applied attribution analysis to isolate the specific contributions of semantic similarity and lexical divergence to overall human quality scores.

The investigation produced four central findings. First, widely used metrics align poorly with human judgments; for example, standard word-overlap metrics like BLEU showed near-zero or even negative correlations with human annotations. Second, reference-free metrics consistently outperformed reference-based versions because typical test candidates share closer lexical proximity to the input text than to a single human reference. Third, existing metrics capture semantic similarity reasonably well but fail entirely to reward lexical divergence, frequently assigning high scores to verbatim copies of the input. Fourth, human evaluation rewards lexical variation only up to a point; beyond a moderate threshold (an edit distance around 0.35), additional wording variation provides no further quality benefit. Incorporating these insights, the proposed metric, ParaScore, achieved the highest human alignment, improving correlation scores over standard embedding metrics across standard and stress-tested benchmarks.

These findings indicate that relying on legacy metrics like BLEU or standard ROUGE creates significant risk in product development and benchmarking by misjudging paraphrase quality and penalizing creative, valid phrasing. Furthermore, the analysis reveals that standard evaluation benchmarks themselves are skewed, as they inadequately penalize copied text and lack natural diversity. Adopting an evaluation approach that incorporates a bounded threshold for wording variation provides a more accurate assessment of natural language generation quality.

Organizations developing or deploying paraphrasing systems should transition away from legacy translation metrics toward evaluation frameworks that explicitly balance semantic preservation with threshold-based lexical divergence, such as ParaScore. Development teams must also build more representative benchmark datasets that include natural lexical variation to avoid overfitting models to superficial word matching.

A current limitation of this work is that the stress-test benchmarks introduced variation by artificially copying input sentences rather than sourcing natural, diverse human phrasing. Nevertheless, confidence in the primary findings remains high given the consistent mathematical and empirical results across multiple metric families and two distinct languages.

Cover for On the Evaluation Metrics for Paraphrase Generation

Abstract

In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts. (2) Most commonly used metrics do not align well with human annotation. Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses. Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation. It possesses the merits of reference-based and reference-free metrics and explicitly models lexical divergence. Based on our analysis and improvements, our proposed reference-based outperforms than reference-free metrics. Experimental results demonstrate that ParaScore significantly outperforms existing metrics. Our codes and toolkit are released in https://github.com/shadowkiller33/ParaScore.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Revisiting Paraphrasing Metrics
  • 2.1 Settings
  • 2.2 Experimental Results
  • 3 Reference-Free vs. Reference-Based
  • 3.1 The Distance Effect
  • 3.2 Average Distance Hypothesis
  • 3.3 Why do Reference-Free Metrics Perform Better on our Benchmarks?
  • 4 Decoupling Semantic Similarity and Lexical Divergence
  • 4.1 Attribution Analysis for Disentanglement
  • 4.2 Performance in Capturing Sim
  • 4.3 Performance in Capturing Div
  • 5 New Metric: ParaScore
  • 5.1 ParaScore
  • 5.2 Experimental Results
  • 6 Related Work
  • 7 Conclusion
  • Limitation
  • Ethical Considerations
  • References
  • A Details of Twitter-Para
  • B Details of BQ-Para
  • C Definition of normalized edit distance
  • D Definition of BERT-iBLEU and iBLEU
  • E A detailed analysis towards BERT-iBLEU

Knowls

  1. Knowl 1 — Paraphrase evaluation benchmarks and protocol

    experimental setup

    The study evaluates paraphrase metrics on English Twitter-Para and Chinese BQ-Para. Twitter-Para contains 761 input sentences, one reference per input, and 7,159 human-rated candidate paraphrases (9.41 candidates per input on average); its quality ratings range from 0 to 1. BQ-Para, introduced by the authors as a Chinese paraphrase-evaluation benchmark, contains 550 inputs, one manually written reference per input, and 5,550 candidates (10 per input), with human quality ratings from 0 to 1. For metrics with tunable hyperparameters, 10% of each benchmark is used for development and tuning, and performance is measured on the remaining 90%. The extended benchmarks add input sentences themselves as candidates: 20% of the candidates are these copies, assigned a human quality score of 0, to test metric robustness when exact copying is present.

  2. Knowl 2 — Existing metrics often correlate weakly with human paraphrase ratings

    empirical result

    The authors compare reference-based metrics, which score a candidate against its reference, with reference-free variants, which score it against the input. The table reports Pearson and Spearman correlations with human quality ratings on Twitter-Para and BQ-Para. METEOR was not evaluated on Chinese BQ-Para because it relies on English WordNet. Reference-free variants perform better for most metric–dataset comparisons, but not for BLEU-4 or BARTScore on Twitter-Para. Correlations are often modest; BLEU-4 is negatively correlated with human ratings on Twitter-Para. Embedding-based BERTScore variants generally correlate better than the n-gram metrics, but their correlations are still limited.

    Metric Twitter Pearson Twitter Spearman BQ Pearson BQ Spearman
    BLEU-4 -0.119 -0.104 0.127 0.144
    BLEU-4.Free -0.113 -0.101 0.109 0.136
    ROUGE-1 0.271 0.276 0.229 0.206
    ROUGE-1.Free 0.292 0.300 0.264 0.232
    ROUGE-2 0.181 0.144 0.226 0.216
    ROUGE-2.Free 0.228 0.189 0.252 0.242
    ROUGE-L 0.249 0.239 0.221 0.204
    ROUGE-L.Free 0.266 0.253 0.260 0.230
    METEOR 0.423 0.418 – –
    METEOR.Free 0.469 0.471 – –
    BERTScore(B) 0.470 0.468 0.332 0.322
    BERTScore(B).Free 0.491 0.488 0.397 0.392
    BERTScore(R) 0.368 0.358 0.387 0.376
    BERTScore(R).Free 0.373 0.361 0.449 0.438
    BARTScore 0.311 0.306 0.241 0.230
    BARTScore.Free 0.295 0.286 0.282 0.263
  3. Knowl 3 — Relative input–reference distances explain when reference-free metrics win

    empirical result

    Let XX be an input, RR its reference, and CC a candidate; lexical distance is measured by normalized edit distance (NED). The authors report that metric–human correlation generally declines as candidates become more distant from the text being scored, with a particularly sharp decline in the most distant quartile for both reference-based and reference-free metrics. They propose the average-distance hypothesis: for a candidate group GG, reference-based metrics tend to do better when the group's average distance to RR is smaller than its average distance to XX, while reference-free metrics tend to do better in the reverse case. In their split, Case I has Dist⁡(C,R)<Dist⁡(C,X)\operatorname{Dist}(C,R)<\operatorname{Dist}(C,X) and Case II has Dist⁡(C,R)>Dist⁡(C,X)\operatorname{Dist}(C,R)>\operatorname{Dist}(C,X). The reported mean correlation difference is reference-free minus reference-based; it reverses sign across the two cases, consistent with the hypothesis.

    Twitter-Para Case I Twitter-Para Case II BQ-Para Case I BQ-Para Case II
    Mean correlation difference -0.110 +0.132 -0.044 +0.079
    Candidate proportion 46.4% 53.6% 15.7% 84.3%
  4. Knowl 4 — Attribution analysis separates semantic and divergence effects

    model/method

    To assess metric behavior on semantic similarity and lexical divergence separately despite having only an overall human score h(X,C)h(X,C), the authors compare candidate pairs (Cj,Ck)(C_j,C_k) for the same input XX. Normalized edit distance is used as a surrogate for divergence, and SimCSE similarity as a surrogate for semantic similarity. For semantic-focused pairs, the edit distances are matched within 0.050.05 while the SimCSE scores differ by at least 0.150.15; the semantic score difference is ΔS=Sim⁡(X,Cj)−Sim⁡(X,Ck)\Delta S=\operatorname{Sim}(X,C_j)-\operatorname{Sim}(X,C_k), and the human-score difference is Δh=h(X,Cj)−h(X,Ck)\Delta h=h(X,C_j)-h(X,C_k). For divergence-focused pairs, the SimCSE scores are matched within 0.050.05 while edit distances differ by at least 0.100.10. Correlation between metric-score differences and Δh\Delta h is then measured within each selected pair set. On the semantic-focused set, the correlation of ΔS\Delta S with Δh\Delta h is higher than in the distance-matched comparison set, supporting the intended separation: Twitter-Para has 583 selected pairs with correlation 0.8050.805, versus 9,158 comparison pairs with 0.3450.345; BQ-Para has 200 selected pairs with 0.6290.629, versus 5,156 comparison pairs with 0.3940.394.

  5. Knowl 5 — Embedding metrics capture semantic similarity better than word overlap

    empirical result

    On semantic-focused candidate pairs—pairs with nearly matched input-to-candidate NED but substantially different SimCSE similarity—the authors correlate each reference-free metric's score difference with the human-score difference. BERTScore and BARTScore show stronger semantic alignment than BLEU-4.Free, while the SimCSE surrogate itself provides a comparison point. The values are Pearson correlations on Twitter-Para and BQ-Para, respectively.

    Metric Twitter-Para BQ-Para
    BLEU-4.Free 0.067 0.372
    ROUGE-1.Free 0.574 0.430
    ROUGE-2.Free 0.400 0.350
    ROUGE-L.Free 0.481 0.388
    METEOR.Free 0.499 –
    BERTScore(B).Free 0.785 0.576
    BARTScore.Free 0.797 0.552
    SimCSE similarity 0.805 0.629
  6. Knowl 6 — Lexical divergence helps only below a distance threshold

    empirical result

    For divergence-focused pairs, the authors hold SimCSE similarity nearly constant (difference at most 0.050.05) while requiring an NED difference of at least 0.100.10. They divide these pairs according to d(j,k)=min⁡(NED⁡(X,Cj),NED⁡(X,Ck))d(j,k)=\min(\operatorname{NED}(X,C_j),\operatorname{NED}(X,C_k)): the lower-distance subset has d(j,k)≤0.35d(j,k)\leq 0.35, and the higher-distance subset has d(j,k)>0.35d(j,k)>0.35. The correlation between the NED difference and human-score difference is substantial in the lower-distance subset but close to zero in the higher-distance subset, indicating that added lexical divergence can correspond to improved quality below the threshold but does not reliably improve quality beyond it. In the lower-distance subset, every tested automatic metric other than NED itself has a negative correlation with human-score differences, consistent with existing metrics failing to reward the relevant divergence.

    Twitter lower Twitter higher BQ lower BQ higher
    Pair count 192 3876 290 6217
    NED-difference correlation with human difference 0.635 0.021 0.655 0.025

    The following are Pearson correlations between each metric-score difference and the human-score difference on the lower-distance subset; NED is included as the divergence reference measure.

    Metric Twitter-Para BQ-Para
    BLEU-4.Free -0.197 -0.075
    ROUGE-1.Free -0.385 -0.334
    ROUGE-2.Free -0.377 -0.308
    ROUGE-L.Free -0.426 -0.514
    METEOR.Free -0.233 –
    BERTScore(B).Free -0.424 -0.347
    BARTScore.Free -0.187 -0.263
    NED 0.635 0.655
  7. Knowl 7 — ParaScore combines best-of-input/reference similarity with thresholded divergence

    model/method

    ParaScore is a reference-based paraphrase metric designed to account for both semantic similarity and lexical divergence. For input XX, reference RR, and candidate CC, it takes the larger of candidate similarity to the input and to the reference, then adds a weighted divergence score:

    ParaScore⁡(X,R,C)=max⁡(Sim⁡(X,C),Sim⁡(R,C))+ω DS⁡(X,C).\operatorname{ParaScore}(X,R,C)=\max(\operatorname{Sim}(X,C),\operatorname{Sim}(R,C))+\omega\,\operatorname{DS}(X,C).

    Here Sim⁡\operatorname{Sim} is instantiated as BERTScore, Dist⁡\operatorname{Dist} as normalized edit distance, d=Dist⁡(X,C)d=\operatorname{Dist}(X,C), ω\omega is a tunable weight, and the threshold γ\gamma is fixed at 0.350.35 in the reported experiments. The divergence score increases with distance up to the threshold and then saturates:

    DS⁡(X,C)={γ,d>γ,dγ+1γ−1,0≤d≤γ.\operatorname{DS}(X,C)= \begin{cases} \gamma, & d>\gamma,\\ d\dfrac{\gamma+1}{\gamma}-1, & 0\leq d\leq\gamma. \end{cases}

    The corresponding reference-free version removes reference similarity and uses Sim⁡(X,C)+ωDS⁡(X,C)\operatorname{Sim}(X,C)+\omega\operatorname{DS}(X,C). The maximum in the reference-based version allows either input or reference similarity to supply the semantic match; the capped divergence term avoids continuing to reward greater distance once it exceeds the chosen threshold.

  8. Knowl 8 — ParaScore has the highest reported correlations on all four benchmarks

    data/table

    The table gives Pearson and Spearman correlations with human ratings for the original English and Chinese benchmarks and for their extended versions, which include copied inputs as zero-rated candidates. ParaScore has the highest reported correlation in each dataset–correlation comparison. Its reference-free version is also strong, but remains below the reference-based score. BERT-iBLEU is comparable to, or below, BERTScore(B).Free in these comparisons.

    Twitter-Para BQ-Para Twitter(Extend) BQ-Para(Extend)
    Metric Pearson Spearman Pearson Spearman Pearson Spearman Pearson Spearman
    BERTScore(B) 0.470 0.468 0.332 0.322 0.427 0.432 0.248 0.267
    BERTScore(R) 0.368 0.358 0.387 0.376 0.334 0.329 0.299 0.317
    BARTScore 0.311 0.306 0.260 0.246 0.280 0.276 0.199 0.206
    iBLEU(0.2) 0.013 0.033 0.155 0.139 0.011 0.032 0.129 0.121
    BERTScore(B).Free 0.491 0.488 0.397 0.392 0.316 0.419 0.230 0.312
    BERT-iBLEU(B,4) 0.488 0.485 0.393 0.383 0.327 0.416 0.221 0.303
    ParaScore 0.522 0.523 0.492 0.489 0.527 0.530 0.510 0.442
    ParaScore.Free 0.492 0.489 0.398 0.393 0.496 0.495 0.487 0.428
  9. Knowl 9 — Ablations support each ParaScore component

    empirical result

    Ablations on the extended benchmarks remove the thresholded divergence formulation, the maximum over input/reference similarity, or the divergence score entirely. Every ablation lowers both reported correlations relative to full ParaScore; removing the divergence score or its thresholded formulation causes the largest drops. This supports the authors' conclusion that the similarity combination, divergence term, and threshold mechanism each contribute to the reported performance.

    Twitter(Extend) Pearson Twitter(Extend) Spearman BQ-Para(Extend) Pearson BQ-Para(Extend) Spearman
    ParaScore 0.527 0.530 0.510 0.442
    ParaScore without threshold 0.358 0.450 0.266 0.333
    ParaScore without maximum 0.496 0.495 0.487 0.428
    ParaScore without divergence score 0.349 0.450 0.249 0.326
  10. Knowl 10 — Benchmark construction does not test natural lexical-divergence variation

    limitation

    The authors identify the lack of a benchmark that reflects lexical divergence in a natural way as a limitation. The extended datasets introduce exact input copies as zero-quality candidates, which tests robustness to copying but adds divergence-related cases heuristically rather than representing a natural range of paraphrases. Consequently, the reported experiments do not establish how ParaScore performs on a benchmark that naturally and faithfully captures the importance of lexical divergence; the authors identify building such a benchmark and evaluating ParaScore on it as future work.

Coverage note — The paper's detailed BERT-iBLEU decomposition is not a separate knowl because its main empirical comparison is included with ParaScore results, while the algebraic diagnostic is metric-specific and secondary.

References

  1. 1.Abdalghani Abujabal, Rishiraj Saha Roy, Mohamed Yahya, and Gerhard Weikum. 2019. Comqa: A community-sourced dataset for complex factoid question answering with paraphrase clusters. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 307–317.
  2. 2.Icek Ajzen and Martin Fishbein. 1975. A bayesian analysis of attribution processes. Psychological bulletin, 82(2):261.
  3. 3.W Thomas Anderson Jr, Eli P Cox III, and David G Fulcher. 1976. Bank selection decisions and market segmentation: Determinant attribute analysis reveals convenience-and sevice-oriented bank customers. Journal of marketing, 40(1):40–45.
  4. 4.Marianna Apidianaki, Guillaume Wisniewski, Anne Cocos, and Chris Callison-Burch. 2018. Automated paraphrase lattice creation for hyter machine translation evaluation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 480–485.
  5. 5.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  6. 6.Colin Bannard and Chris Callison-Burch. 2005. Paraphrasing with bilingual parallel corpora. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 597–604.
  7. 7.Rahul Bhagat and Eduard Hovy. 2013. What is a paraphrase? Computational Linguistics, 39(3):463–472.
  8. 8.Deng Cai, Yan Wang, Huayang Li, Wai Lam, and Lemao Liu. 2021. Neural machine translation with monolingual translation memory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7307–7318.
  9. 9.Chris Callison-Burch. 2008. Syntactic constraints on paraphrases extracted from parallel corpora. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 196–205.
  10. 10.Ruisheng Cao, Su Zhu, Chenyu Yang, Chen Liu, Rao Ma, Yanbin Zhao, Lu Chen, and Kai Yu. 2020. Unsupervised dual paraphrasing for two-stage semantic parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6806–6817.
  11. 11.Zhangming Chan, Lemao Liu, Juntao Li, Haisong Zhang, Dongyan Zhao, Shuming Shi, and Rui Yan. 2021. Enhancing the open-domain dialogue evaluation in latent space. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021.
  12. 12.David Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 190–200.
  13. 13.Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang. 2018. The bq corpus: A large-scale domain-specific chinese corpus for sentence semantic equivalence identification. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 4946–4951.
  14. 14.Wang Chen, Piji Li, and Irwin King. 2021. A training-free and reference-free summarization evaluation metric via centrality-weighted relevance and self-referenced redundancy. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics.
  15. 15.Leshem Choshen and Omri Abend. 2018. Automatic metric validation for grammatical error correction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1372–1382.
  16. 16.Trevor Cohn, Chris Callison-Burch, and Mirella Lapata. 2008. Constructing corpora for the development and evaluation of paraphrase systems. Computational Linguistics, 34(4):597–614.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  18. 18.Li Dong, Jonathan Mallinson, Siva Reddy, and Mirella Lapata. 2017. Learning to paraphrase for question answering. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 875–886.
  19. 19.Steven Y Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021. A survey of data augmentation approaches for nlp. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 968–988.
  20. 20.Markus Freitag, David Grangier, and Isaac Caswell. 2020. Bleu might be guilty but references are not innocent. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 61–71.
  21. 21.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation.
  22. 22.Wee Chung Gan and Hwee Tou Ng. 2019. Improving the robustness of question answering systems to question paraphrasing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6065–6075.
  23. 23.Jun Gao, Wei Bi, Ruifeng Xu, and Shuming Shi. 2021a. Ream: An enhancement approach to reference-based evaluation metrics for open-domain dialog generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics: Findings.
  24. 24.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021b. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910.
  25. 25.Lila R Gleitman and Henry Gleitman. 1970. Phrase and paraphrase: Some innovative uses of language.
  26. 26.Tanya Goyal and Greg Durrett. 2020. Neural syntactic preordering for controlled paraphrase generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 238–252.
  27. 27.Qiuxiang He, Guoping Huang, Qu Cui, Li Li, and Lemao Liu. 2021. Fast and accurate neural machine translation with translation memory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3170–3180.
  28. 28.Chaitra Hegde and Shrikumar Patil. 2020. Unsupervised paraphrase generation using pre-trained language models. arXiv preprint arXiv:2006.05477.
  29. 29.Jonathan Herzig and Jonathan Berant. 2019. Don’t paraphrase, detect! rapid and effective data collection for semantic parsing. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3810–3820.
  30. 30.Kuan-Hao Huang and Kai-Wei Chang. 2021. Generating syntactically controlled paraphrases without using annotated parallel pairs. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1022–1033.
  31. 31.Tomoyuki Kajiwara. 2019. Negative lexically constrained decoding for paraphrase generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6047–6052.
  32. 32.Ashutosh Kumar, Kabir Ahuja, Raghuram Vadapalli, and Partha Talukdar. 2020. Syntax-guided controlled generation of paraphrases. Transactions of the Association for Computational Linguistics, 8:330–345.
  33. 33.Ashutosh Kumar, Satwik Bhattamishra, Manik Bhandari, and Partha Talukdar. 2019. Submodular optimization-based diverse paraphrasing and its effectiveness in data augmentation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3609–3619.
  34. 34.Wuwei Lan and Wei Xu. 2018. Neural network models for paraphrase identification, semantic textual similarity, natural language inference, and question answering. In Proceedings of the 27th International Conference on Computational Linguistics, pages 3890–3902.
  35. 35.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  36. 36.Xianggen Liu, Lili Mou, Fandong Meng, Hao Zhou, Jie Zhou, and Sen Song. 2020. Unsupervised paraphrasing by simulated annealing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 302–312.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  38. 38.Nitin Madnani, Joel Tetreault, and Martin Chodorow. 2012. Re-examining machine translation metrics for paraphrase identification. In Proceedings of the 2012 conference of the north american chapter of the association for computational linguistics: Human language technologies, pages 182–190.
  39. 39.George A Miller. 1995. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41.
  40. 40.Tong Niu, Semih Yavuz, Yingbo Zhou, Nitish Shirish Keskar, Huan Wang, and Caiming Xiong. 2021. Unsupervised paraphrasing with pretrained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5136–5150.
  41. 41.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  42. 42.Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André FT Martins, and Alon Lavie. 2021. Are references really needed? unbabel-ist 2021 submission for the metrics shared task. In Proceedings of the Sixth Conference on Machine Translation, pages 1030–1040.
  43. 43.Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. Bleurt: Learning robust metrics for text generation. ACL.
  44. 44.Xiaoyu Shen, Hui Su, Wenjie Li, and Dietrich Klakow. 2018. Nexus network: Connecting the preceding and the following in dialogue generation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4316–4327.
  45. 45.Shuming Shi, Enbo Zhao, Duyu Tang, Yan Wang, Piji Li, Wei Bi, Haiyun Jiang, Guoping Huang, Leyang Cui, Xinting Huang, et al. 2022. Effidit: Your ai writing assistant. arXiv preprint arXiv:2208.01815.
  46. 46.Raphael Shu, Hideki Nakayama, and Kyunghyun Cho. 2019. Generating diverse translations with sentence codes. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1823–1827.
  47. 47.AB Siddique, Samet Oymak, and Vagelis Hristidis. 2020. Unsupervised paraphrasing via deep reinforcement learning. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1800–1809.
  48. 48.Jiao Sun, Xuezhe Ma, and Nanyun Peng. 2021. Aesop: Paraphrase generation with adaptive syntactic control. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5176–5189.
  49. 49.Clifford H Wagner. 1982. Simpson’s paradox in real life. The American Statistician, 36(1):46–48.
  50. 50.Shan Wu, Bo Chen, Chunlei Xin, Xianpei Han, Le Sun, Weipeng Zhang, Jiansong Chen, Fan Yang, and Xunliang Cai. 2021. From paraphrasing to semantic parsing: Unsupervised semantic parsing via synchronous semantic decoding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5110–5121.
  51. 51.Jiannan Xiang, Huayang Li, Yahui Liu, Lemao Liu, Guoping Huang, Defu Lian, and Shuming Shi. 2022. Investigating data variance in evaluations of automatic machine translation metrics. In Findings of the Association for Computational Linguistics: ACL 2022.
  52. 52.Jiannan Xiang, Yahui Liu, Deng Cai, Huayang Li, Defu Lian, and Lemao Liu. 2021. Assessing dialogue systems with distribution distances. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics: Findings.
  53. 53.Wei Xu, Chris Callison-Burch, and William B Dolan. 2015. Semeval-2015 task 1: Paraphrase and semantic similarity in twitter (pit). In Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015), pages 1–11.
  54. 54.Wei Xu, Alan Ritter, Chris Callison-Burch, William B Dolan, and Yangfeng Ji. 2014. Extracting lexically divergent paraphrases from twitter. Transactions of the Association for Computational Linguistics, 2:435–448.
  55. 55.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34.
  56. 56.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. ICLR.
  57. 57.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019. Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. arXiv preprint arXiv:1909.02622.
  58. 58.Jianing Zhou and Suma Bhat. 2021. Paraphrase generation: A survey of the state of the art. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5075–5086.

Citation

MLA
Shen, L., et al. “On the Evaluation Metrics for Paraphrase Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3178–90, https://doi.org/10.18653/v1/2022.emnlp-main.208.
APA
Shen, L., Liu, L., Jiang, H., & Shi, S. (2022). On the Evaluation Metrics for Paraphrase Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3178–3190. https://doi.org/10.18653/v1/2022.emnlp-main.208
Chicago
Shen, L., L. Liu, H. Jiang, and S. Shi. 2022. “On the Evaluation Metrics for Paraphrase Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3178–90. https://doi.org/10.18653/v1/2022.emnlp-main.208.
Harvard
Shen, L. et al. (2022) “On the Evaluation Metrics for Paraphrase Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3178–3190. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.208.
Vancouver
1. Shen L, Liu L, Jiang H, Shi S (2022) On the Evaluation Metrics for Paraphrase Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3178–3190

BibTeX

@inproceedings{shen-etal-2022-evaluation,
    title = "On the Evaluation Metrics for Paraphrase Generation",
    author = "Shen, Lingfeng  and
      Liu, Lemao  and
      Jiang, Haiyun  and
      Shi, Shuming",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.208/",
    doi = "10.18653/v1/2022.emnlp-main.208",
    pages = "3178--3190"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/