GPTScore: Evaluate as You Desire

Jinlan FuSee-Kiong NgZhengbao JiangPengfei Liu

article2024NAACL504 citations

Proposes GPTScore, a training-free evaluation framework that leverages the conditional generation probabilities of large language models to assess generated text across custom criteria defined purely through natural language instructions.

Listen

As generative artificial intelligence rapidly evolves to produce high-quality text, techniques for evaluating these outputs have failed to keep pace. Conventional automated evaluation metrics typically focus on narrow dimensions, require costly model training or human data labeling, and cannot easily adapt to customized assessment criteria required by practitioners. To overcome these constraints, the article introduces GPTScore, a text evaluation framework that scores generated content without requiring specialized model training by leveraging the conditional generation probabilities of pre-trained language models guided by natural language instructions.

The article's primary objective is to demonstrate that pre-trained language models can achieve customizable, multi-faceted, and training-free text evaluation across diverse applications. To assess this, the authors evaluated 19 pre-trained language models ranging in size from 80 million to 175 billion parameters across four major natural language generation tasks: text summarization, data-to-text generation, dialogue response generation, and machine translation. The evaluation benchmark encompassed 37 datasets across 22 distinct quality aspects, measuring automated scores against human judgments using statistical correlation benchmarks.

The analysis established several key findings. First, instructing models with explicit task and quality aspect definitions significantly improved evaluation accuracy compared to uninstructed baselines, with few-shot demonstration examples yielding further gains across most tasks. Second, training-free GPTScore models routinely matched or outperformed established supervised evaluators that underwent task-specific fine-tuning. Third, in dialogue response evaluation, the older 175-billion-parameter text-davinci-001 model drastically outperformed the human-feedback-tuned text-davinci-003, achieving an average correlation improvement of 40.8 points on turn-level assessments. Finally, enriching evaluation prompts by combining correlated quality dimensions enabled medium-sized models, such as the 6.7-billion-parameter GPT-3 Curie, to match or exceed the performance of 175-billion-parameter models.

These results demonstrate that organizations can implement highly customized, multi-dimensional quality auditing for generative text without incurring the substantial computational and annotation costs of training dedicated evaluators. However, the findings also highlight that model size and recent human-feedback tuning do not guarantee superior evaluation performance, meaning practitioners must carefully validate evaluator configurations rather than assuming newer or larger models will perform best. For production deployments, smaller and properly prompted models present a cost-effective alternative to very large commercial models.

Decision-makers aiming to adopt this framework should define clear, explicit instructions for required quality criteria and test prompt compositions that integrate related dimensions before scaling. Certain limitations must be considered: newer proprietary models such as GPT-4 were not evaluated, the internal behavior of proprietary models remains opaque, and language-model evaluators carry inherent risks of unobservable bias. The evidence strongly supports the feasibility of the framework, but organizations should conduct pilot validations and safety audits prior to high-stakes deployment.

Cover for GPTScore: Evaluate as You Desire

Abstract

Generative Artificial Intelligence (AI) has enabled the development of sophisticated models that are capable of producing high-caliber text, images, and other outputs through the utilization of large pre-trained models. Nevertheless, assessing the quality of the generation is an even more arduous task than the generation itself, and this issue has not been given adequate consideration recently. This paper proposes a novel evaluation framework, GPTSCORE, which utilizes the emergent abilities (e.g., in-context learning, zero-shot instruction) of generative pre-trained models to score generated texts. There are 19 pre-trained models explored in this paper, ranging in size from 80M (e.g., Flan-T5-small) to 175B (e.g., GPT3). Experimental results on four text generation tasks, 22 evaluation aspects, and corresponding 37 datasets demonstrate that this approach can effectively allow us to achieve what one desires to evaluate for texts simply by natural language instructions. This nature helps us overcome several long-standing challenges in text evaluation—how to achieve customized, multi-faceted evaluation without model training. We make our code publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Generative Pretraining Score (GPTScore)
  • 4 Experimental Settings
  • 4.1 Meta Evaluation
  • 4.2 Tasks, Datasets, and Aspects
  • 4.3 Scoring Models
  • 4.4 Scoring Dimension
  • 4.5 Evaluation Dataset Construction
  • 5 Experiment Results
  • 5.1 Text Summarization
  • 5.2 Data to Text
  • 5.3 Dialogue Response Generation
  • 6 Ablation Study
  • 6.1 Effectiveness of Demonstration
  • 6.2 Partial Order of Evaluation Aspect
  • 7 Conclusion
  • 8 Limitations
  • Acknowledgements
  • References
  • A Machine Translation
  • B Evaluation Strategy
  • C Metric Comparison
  • D Tasks, Datasets, and Aspects
  • E Ablation Study
  • E.1 Effectiveness of Demonstration
  • E.2 Partial Order of Evaluation Aspect
  • F Prompt Design
  • G Experiment Results

Knowls

  1. Knowl 1 — GPTSCORE: Training-Free Multi-Aspect Text Evaluation via Generative Pre-training

    model/method

    GPTSCORE is a training-free text evaluation framework that evaluates the quality of generated text by computing the conditional generation probability from pre-trained generative language models guided by natural language instructions and optional exemplar demonstrations.

    Rather than training task-specific or aspect-specific neural evaluators on human rating datasets, GPTSCORE constructs an evaluation protocol using a prompt template T(d,a,S)T(d, a, S) composed of three elements:

    1. Task specification (dd): Describes how the text is generated (e.g., generating a summary for a given document or producing a dialogue response).
    2. Aspect definition (aa): Defines the explicit quality dimension to be evaluated (e.g., fluency, coherence, factuality, informativeness, or custom user-defined criteria).
    3. Context text (SS): Contains input context information, such as the source document, reference text, or dialogue history.

    Additionally, KK few-shot exemplar demonstrations can be prepended to the prompt to provide in-context guidance. The evaluated candidate hypothesis hh is evaluated directly by calculating its conditional generation probability under the model parameterized by θ\theta.

  2. Knowl 2 — Mathematical Formulation of GPTSCORE

    equation

    Let h=(h1,h2,…,hm)h = (h_1, h_2, \dots, h_m) be a candidate text hypothesis consisting of mm tokens. Let dd denote the task description, aa denote the aspect definition, and SS denote the context text (such as source text or gold reference text). Let T(d,a,S)T(d, a, S) denote the prompt template structuring the task, aspect, and context into an evaluation protocol. For a pre-trained generative language model parameterized by θ\theta, GPTSCORE is defined as the weighted conditional log-likelihood:

    GPTScore(h∣d,a,S)=∑t=1mwtlog⁡p(ht∣h<t,T(d,a,S),θ)\text{GPTScore}(h \mid d, a, S) = \sum_{t=1}^m w_t \log p(h_t \mid h_{<t}, T(d, a, S), \theta)

    where h<t=(h1,…,ht−1)h_{<t} = (h_1, \dots, h_{t-1}) is the token prefix prior to index tt, and wt∈Rw_t \in \mathbb{R} is the weight assigned to token tt. In standard GPTSCORE implementations, all token weights are set equally (wt=1w_t = 1). Higher conditional log-likelihood values correspond to candidate texts that more closely fulfill the quality criteria specified in the instruction prompt.

  3. Knowl 3 — Scoring Dimensions and Text Input Directions in GPTSCORE

    definition

    GPTSCORE uses different scoring directions to structure the input context and evaluated output depending on how human judgments are gathered for a given task:

    1. Source-to-Hypothesis (src→hypo\text{src} \to \text{hypo}): Computes the conditional probability p(hypo∣src)p(\text{hypo} \mid \text{src}). This direction is used when human annotators judge candidate texts directly against the source input, such as dialogue response generation (evaluating a response conditioned on conversation history) and summarization aspects (evaluating consistency, coherence, and informativeness against the source document).

    2. Reference-to-Hypothesis (ref→hypo\text{ref} \to \text{hypo} or ref↔hypo\text{ref} \leftrightarrow \text{hypo}): Computes the conditional probability p(hypo∣ref)p(\text{hypo} \mid \text{ref}) (or bidirectional evaluation). This direction is used when the source is not in standard unstructured natural language text (such as structured table inputs in data-to-text generation) or when the source is in a different language from the hypothesis (such as machine translation evaluated with monolingual language models).

  4. Knowl 4 — Synergistic Aspect Definition Composition

    empirical result

    In GPTSCORE, expanding the aspect definition in the prompt by composing it with definitions of correlated quality dimensions substantially improves correlation with human judgments, allowing smaller models to surpass larger models evaluated on single-aspect prompts.

    On the FED Turn-level dialogue response benchmark evaluating the aspect "interesting" (INT), incrementally combining the target definition with correlated aspect definitions in descending order of human correlation produces steady improvements using GPT-3 Curie (6.7B parameters):

    • 1 aspect (INT: "Is this response interesting to the conversation?"): Spearman correlation ρ=36.9\rho = 36.9
    • 2 aspects (INT + ENG [Engaging]): ρ=40.7\rho = 40.7
    • 3 aspects (INT + ENG + SPE [Specific]): ρ=48.6\rho = 48.6
    • 4 aspects (INT + ENG + SPE + COR [Correct]): ρ=50.0\rho = 50.0
    • 5 aspects (INT + ENG + SPE + COR + REL [Relevant]): ρ=51.3\rho = 51.3
    • 7 aspects (adding UND [Understandable] and SEM [Semantically appropriate]): ρ=51.4\rho = 51.4

    Through aspect composition, the 6.7B-parameter GPT-3 Curie model reaches ρ=51.4\rho = 51.4, outperforming the 175B-parameter GPT-3 Davinci-001 model evaluated on the single INT aspect definition (ρ=50.1\rho = 50.1).

  5. Knowl 5 — Text Summarization Meta-Evaluation Across Model Backbones

    data/table

    Meta-evaluation on the SummEval benchmark measures sample-level Spearman correlation (ρ\rho) between metric scores and human judgments across four aspects: consistency (CON), fluency (FLU), relevance (REL), and coherence (COH). GPTSCORE is evaluated in vanilla (VAL, unprompted) and instruction (IST, with task description and aspect definition) settings across 19 model variants covering GPT-3, GPT-2, OPT, and Flan-T5 backbones.

    Model CON FLU REL COH
    VAL IST VAL IST VAL IST VAL IST
    ROUGE-1 20.8 - 14.8 - 26.2 - 14.1 -
    ROUGE-2 17.2 - 12.0 - 17.4 - 9.1 -
    ROUGE-L 19.8 - 17.6 - 24.7 - 12.9 -
    BERTScore 19.7 - 23.7 - 34.7 - 25.9 -
    MoverScore 18.0 - 15.7 - 24.8 - 11.5 -
    PRISM 29.9 - 26.1 - 25.2 - 26.5 -
    BARTScore 30.8 - 24.6 - 28.9 - 29.7 -
    BARTScore+CNN 35.8 - 38.1 - 35.9 - 42.5 -
    BARTScore+CNN+Para 37.0 - 40.5 - 33.9 - 42.5 -
    GPT3-text-ada-001 (350M) 39.7 40.5 36.1 35.9 28.2 27.6 39.3 39.8
    GPT3-text-babbage-001 (1.3B) 41.0 41.4 37.1 39.1 32.0 33.4 42.7 45.2
    GPT3-text-curie-001 (6.7B) 44.6 45.1 38.9 39.5 31.6 33.2 41.3 40.8
    GPT3-text-davinci-001 (175B) 46.6 47.5 40.5 41.0 32.4 34.3 40.0 40.1
    GPT3-text-davinci-003 (175B) 45.2 44.9 41.1 40.3 36.3 38.1 43.7 43.4
    Flan-T5-small (80M) 37.0 38.0 35.6 34.7 27.3 28.0 35.0 35.4
    Flan-T5-Large (770M) 41.0 42.5 39.3 41.6 31.2 35.3 42.3 45.1
    Flan-T5-XL (3B) 41.0 43.6 39.7 42.1 31.4 34.4 42.8 47.0
    Flan-T5-XXL (11B) 43.7 43.8 39.8 42.4 32.8 34.3 42.1 45.6
    Average (all 19 variants) 40.4 41.4 35.8 37.2 31.3 32.2 38.0 40.2

    Equipping pre-trained models with explicit natural language instructions significantly improves average correlation across all aspects. Furthermore, training-free Flan-T5-XL (3B) and Flan-T5-XXL (11B) with instructions outperform the fine-tuned baseline BARTScore+CNN+Para across all four aspects.

  6. Knowl 6 — Dialogue Response Meta-Evaluation and Davinci-001 versus Davinci-003 Disparity

    empirical result

    On the FED dialogue evaluation benchmark, automated metrics are meta-evaluated using dataset-level Spearman correlation across 11 dialogue-level aspects and 8 turn-level aspects against human annotations.

    Evaluation Level Baselines GPTScore (GPT-3 Backbones)
    BARTScore FED DynaEval text-ada-001 text-curie-001 text-davinci-001 text-davinci-003
    FED Dialogue-Level (Avg) 5.8 20.4 39.8 25.3 28.6 54.3 13.5
    FED Turn-Level (Avg) 12.8 11.9 25.6 16.1 25.4 38.3 32.8

    Fine-tuned dialogue baselines such as DynaEval (39.8 dialogue-level, 25.6 turn-level) and FED (20.4 dialogue-level, 11.9 turn-level) outperform standard BARTScore. However, GPTSCORE with GPT3-text-davinci-001 (175B) achieves 54.3 on dialogue-level and 38.3 on turn-level, outperforming all fine-tuned dialogue evaluators.

    Despite sharing the same 175B scale, GPT3-text-davinci-003 (aligned via reinforcement learning from human feedback / RLHF) performs substantially worse than GPT3-text-davinci-001 (13.5 vs. 54.3 on dialogue-level, and 32.8 vs. 38.3 on turn-level), demonstrating that RLHF tuning can impair token-level generation probabilities for text quality scoring.

  7. Knowl 7 — Data-to-Text Meta-Evaluation with Instruction and In-Context Demonstrations

    empirical result

    In data-to-text evaluation across BAGEL and SFRES datasets covering informativeness (INF), naturalness (NAT), and fluency (FLU), GPTSCORE is evaluated in three configurations:

    1. Vanilla (VAL): Evaluates text generation without instructions or demonstrations.
    2. Instruction (IST): Adds natural language task specifications and aspect definitions.
    3. Instruction + Demonstration (IDM): Combines instructions with few-shot demonstration exemplars.

    Average Spearman correlations across all 19 model variants show clear progression:

    • BAGEL:
      • INF: VAL=39.1→IST=40.6→IDM=40.3\text{VAL} = 39.1 \to \text{IST} = 40.6 \to \text{IDM} = 40.3
      • NAT: VAL=27.7→IST=29.8→IDM=33.2\text{VAL} = 27.7 \to \text{IST} = 29.8 \to \text{IDM} = 33.2
      • FLU: VAL=35.8→IST=37.6→IDM=41.6\text{VAL} = 35.8 \to \text{IST} = 37.6 \to \text{IDM} = 41.6
    • SFRES:
      • INF: VAL=25.5→IST=24.7→IDM=24.0\text{VAL} = 25.5 \to \text{IST} = 24.7 \to \text{IDM} = 24.0
      • NAT: VAL=29.1→IST=31.7→IDM=34.2\text{VAL} = 29.1 \to \text{IST} = 31.7 \to \text{IDM} = 34.2
      • FLU: VAL=23.6→IST=26.8→IDM=28.2\text{VAL} = 23.6 \to \text{IST} = 26.8 \to \text{IDM} = 28.2

    On BAGEL, GPT3-text-curie-001 (6.7B) with IDM reaches INF=47.5\text{INF} = 47.5, NAT=39.9\text{NAT} = 39.9, and FLU=44.2\text{FLU} = 44.2, outperforming both GPT3-text-davinci-003 and the supervised fine-tuned baseline BARTScore+CNN+Para (which obtains INF=39.2\text{INF} = 39.2, NAT=31.0\text{NAT} = 31.0, and FLU=44.9\text{FLU} = 44.9).

  8. Knowl 8 — Demonstration Sample Size Scaling Dynamics in GPTSCORE

    empirical result

    On the MQM-2020 machine translation benchmark evaluating accuracy (ACC), fluency (FLU), and multidimensional quality metrics (MQM), varying the number of few-shot demonstration exemplars K∈{0,1,2,4,8,12}K \in \{0, 1, 2, 4, 8, 12\} reveals distinct scaling characteristics across GPT-3 backbones (from 350M to 175B parameters):

    1. Performance increases with demonstration count: Adding demonstrations generally boosts Spearman correlation with human judgments across all quality aspects.
    2. Diminishing returns beyond K=4K = 4: Correlation gains plateau at K≥4K \ge 4. For ACC with GPT3-text-davinci-001, ρ\rho increases from 26.9 at K=0K = 0 to 30.3 at K=4K = 4, 31.2 at K=8K = 8, and 31.7 at K=12K = 12.
    3. Sensitivity and degradation in small models at K=1K = 1: When using only a single exemplar (K=1K = 1), small models such as GPT3-text-ada-001 (350M) suffer performance drops (ACC decreases from 23.7 at K=0K = 0 to 22.5 at K=1K = 1; FLU decreases from 6.3 at K=0K = 0 to 4.9 at K=1K = 1) due to overfitting to non-representative features of a single sample.
  9. Knowl 9 — Sample-Level versus Dataset-Level Meta-Evaluation Metrics

    definition

    Meta-evaluation calculates the correlation between automated metric scores and human judgment scores. Let sis_i for i∈{1,…,n}i \in \{1, \dots, n\} denote source instances, and let hi,jh_{i,j} for j∈{1,…,J}j \in \{1, \dots, J\} denote candidate texts produced by system jj. Let fauto(hi,j)f_{\text{auto}}(h_{i,j}) denote the automatic metric score and fhuman(hi,j)f_{\text{human}}(h_{i,j}) denote human rating. Two aggregation strategies are defined:

    1. Sample-level correlation (Ffauto,fhumansampleF_{f_{\text{auto}}, f_{\text{human}}}^{\text{sample}}): Correlation function gg (Spearman ρ\rho or Pearson rr) is calculated across system outputs separately for each sample ii, then averaged across all nn samples:
    Ffauto,fhumansample=1n∑i=1ng([fauto(hi,1),…,fauto(hi,J)],[fhuman(hi,1),…,fhuman(hi,J)])F_{f_{\text{auto}}, f_{\text{human}}}^{\text{sample}} = \frac{1}{n} \sum_{i=1}^n g\Big(\big[f_{\text{auto}}(h_{i,1}), \dots, f_{\text{auto}}(h_{i,J})\big], \big[f_{\text{human}}(h_{i,1}), \dots, f_{\text{human}}(h_{i,J})\big]\Big)

    This strategy is applied to text summarization, data-to-text, and machine translation benchmarks.

    1. Dataset-level correlation (Ffauto,fhumandataF_{f_{\text{auto}}, f_{\text{human}}}^{\text{data}}): Evaluates correlation globally across all n×Jn \times J system outputs simultaneously:
    Ffauto,fhumandata=g([fauto(h1,1),…,fauto(hn,J)],[fhuman(h1,1),…,fhuman(hn,J)])F_{f_{\text{auto}}, f_{\text{human}}}^{\text{data}} = g\Big(\big[f_{\text{auto}}(h_{1,1}), \dots, f_{\text{auto}}(h_{n,J})\big], \big[f_{\text{human}}(h_{1,1}), \dots, f_{\text{human}}(h_{n,J})\big]\Big)

    This strategy is applied to dialogue response generation benchmarks.

  10. Knowl 10 — Limitations of GPTSCORE and Generative Language Model Evaluators

    limitation

    The GPTSCORE framework and generative LLM-based evaluation exhibit several key limitations:

    1. Model generation recency: The pre-trained language models evaluated are limited to models released prior to or including the GPT-3.5 series; newer chat-aligned models (such as ChatGPT and GPT-4) are not assessed.
    2. Negative impact of RLHF alignment: GPT3-text-davinci-003, which underwent reinforcement learning from human feedback, underperforms the instruction-tuned GPT3-text-davinci-001 in many evaluation settings, indicating an unresolved conflict between RLHF alignment and conditional log-likelihood scoring calibration.
    3. Model opacity: Proprietary commercial LLMs have undisclosed architectures, training corpora, and alignment procedures, limiting explanatory insight into metric behavior and failure modes.
    4. Bias and evaluation safety: Pre-trained generative models can introduce social bias or preference artifacts acquired during pre-training and feedback tuning into text quality scores.
    5. Task scope: API costs restrict empirical investigation to four standard NLG tasks (summarization, dialogue, data-to-text, and machine translation), omitting open-ended long-form generation tasks such as creative story generation.

Coverage note — None was omitted; the knowls cover the GPTSCORE formulation, scoring dimension variants, aspect definition composition, experimental evaluations across summarization, dialogue, data-to-text, and machine translation, demonstration scaling ablation, meta-evaluation aggregation metrics, and paper limitations.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020. Towards a human-like open-domain chatbot. CoRR, abs/2001.09977.
  2. 2.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating evaluation in text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 9347–9359. Association for Computational Linguistics.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  4. 4.Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual error correction for abstractive summarization models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6251–6258. Association for Computational Linguistics.
  5. 5.Louis Castricato, Alexander Havrilla, Shahbuland Matiana, Michael Pieler, Anbang Ye, Ian Yang, Spencer Frazier, and Mark O. Riedl. 2022. Robust preference learning for storytelling via contrastive reinforcement learning. CoRR, abs/2210.07792.
  6. 6.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  7. 7.Esin Durmus, He He, and Mona T. Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 5055–5070. Association for Computational Linguistics.
  8. 8.Markus Freitag, George F. Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. Experts, errors, and context: A large-scale study of human evaluation for machine translation. CoRR, abs/2104.14478.
  9. 9.Jinlan Fu, See-Kiong Ng, and Pengfei Liu. 2022. Polyglot prompt: Multilingual multitask promptraining. arXiv preprint arXiv:2204.14264.
  10. 10.Sarik Ghazarian, Nuan Wen, Aram Galstyan, and Nanyun Peng. 2022. DEAM: dialogue coherence evaluation using amr-based semantic manipulations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 771–785. Association for Computational Linguistics.
  11. 11.Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 708–719. Association for Computational Linguistics.
  12. 12.Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 1693–1701.
  13. 13.J. Edward Hu, Abhinav Singh, Nils Holzenberger, Matt Post, and Benjamin Van Durme. 2019. Large-scale, diverse, paraphrastic bitexts via sampling and clustering. In Proceedings of the 23rd Conference on Computational Natural Language Learning, CoNLL 2019, Hong Kong, China, November 3-4, 2019, pages 44–54. Association for Computational Linguistics.
  14. 14.Matt J. Kusner, Yu Sun, Nicholas I. Kolkin, and Kilian Q. Weinberger. 2015. From word embeddings to document distances. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pages 957–966. JMLR.org.
  15. 15.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  16. 16.Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou. 2021. Conversations are not flat: Modeling the dynamic information flow across dialogue utterances. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 128–138. Association for Computational Linguistics.
  17. 17.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  18. 18.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pretrain, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  19. 19.Yixin Liu, Alexander R Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, et al. 2022. Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation. arXiv preprint arXiv:2212.07981.
  20. 20.François Mairesse, Milica Gasic, Filip Jurcícek, Simon Keizer, Blaise Thomson, Kai Yu, and Steve J. Young. 2010. Phrase-based statistical language generation using graphical models and active learning. In ACL 2010, Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, July 11-16, 2010, Uppsala, Sweden, pages 1552–1561. The Association for Computer Linguistics.
  21. 21.Shahbuland Matiana, J. R. Smith, Ryan Teehan, Louis Castricato, Stella Biderman, Leo Gao, and Spencer Frazier. 2021. Cut the CARP: fishing for zero-shot story evaluation. CoRR, abs/2110.03111.
  22. 22.Shikib Mehri and Maxine Eskénazi. 2020a. Unsupervised evaluation of interactive dialog with dialogpt. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, SIGdial 2020, 1st virtual meeting, July 1-3, 2020, pages 225–235. Association for Computational Linguistics.
  23. 23.Shikib Mehri and Maxine Eskénazi. 2020b. USR: an unsupervised and reference free evaluation metric for dialog generation. CoRR, abs/2005.00456.
  24. 24.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? CoRR, abs/2202.12837.
  25. 25.Mavuto M Mukaka. 2012. A guide to appropriate use of correlation coefficient in medical research. Malawi medical journal, 24(3):69–71.
  26. 26.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  27. 27.Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, and Kewei Tu. 2020. Towards holistic and automatic evaluation of open-domain dialogue generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 3619–3629. Association for Computational Linguistics.
  28. 28.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA, pages 311–318. ACL.
  29. 29.Maja Popovic. 2015. chrf: character n-gram f-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, WMT@EMNLP 2015, 17-18 September 2015, Lisbon, Portugal, pages 392–395. The Association for Computer Linguistics.
  30. 30.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  31. 31.Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. CoRR, abs/2009.09025.
  32. 32.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  33. 33.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. CoRR, abs/2211.05100.
  34. 34.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. Questeval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6594–6604. Association for Computational Linguistics.
  35. 35.Thibault Sellam, Dipanjan Das, and Ankur P. Parikh. 2020. BLEURT: learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7881–7892. Association for Computational Linguistics.
  36. 36.Team Sequoia. 2022. Generative ai: A creative new world. https://www.sequoiacap.com/article/generative-ai-a-creative-new-world/.
  37. 37.Brian Thompson and Matt Post. 2020. Automatic machine translation evaluation in many languages via zero-shot paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 90–121. Association for Computational Linguistics.
  38. 38.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020a. Asking and answering questions to evaluate the factual consistency of summaries. CoRR, abs/2004.04228.
  39. 39.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020b. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 5008–5020. Association for Computational Linguistics.
  40. 40.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926.
  41. 41.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. URL https://arxiv.org/abs/2204.07705.
  42. 42.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  43. 43.Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Peihao Su, David Vandyke, and Steve J. Young. 2015. Semantically conditioned lstm-based natural language generation for spoken dialogue systems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pages 1711–1721. The Association for Computational Linguistics.
  44. 44.Wenda Xu, Xian Qian, Mingxuan Wang, Lei Li, and William Yang Wang. 2022a. Sescore2: Retrieval augmented pretraining for text generation evaluation. CoRR, abs/2212.09305.
  45. 45.Wenda Xu, Yi-Lin Tuan, Yujie Lu, Michael Saxon, Lei Li, and William Yang Wang. 2022b. Not all errors are equal: Learning text generation metrics using stratified error synthesis. In Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 6559–6574. Association for Computational Linguistics.
  46. 46.Zheng Ye, Liucun Lu, Lishan Huang, Liang Lin, and Xiaodan Liang. 2021. Towards quantifiable dialogue coherence evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2718–2729. Association for Computational Linguistics.
  47. 47.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34:27263–27277.
  48. 48.Jerrold H Zar. 2005. Spearman rank correlation. Encyclopedia of biostatistics, 7.
  49. 49.Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021. Dynaeval: Unifying turn and dialogue level evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5676–5689. Association for Computational Linguistics.
  50. 50.Chen Zhang, Luis Fernando D’Haro, Qiquan Zhang, Thomas Friedrichs, and Haizhou Li. 2022a. Finedeval: Fine-grained automatic dialogue-level evaluation. CoRR, abs/2210.13832.
  51. 51.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022b. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  52. 52.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  53. 53.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 563–578. Association for Computational Linguistics.
  54. 54.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation. CoRR, abs/2210.07197.

Citation

MLA
Fu, J., et al. “GPTScore: Evaluate as You Desire”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 6556–76, https://doi.org/10.18653/v1/2024.naacl-long.365.
APA
Fu, J., Ng, S. K., Jiang, Z., & Liu, P. (2024). GPTScore: Evaluate as You Desire. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 6556–6576. https://doi.org/10.18653/v1/2024.naacl-long.365
Chicago
Fu, J., S. K. Ng, Z. Jiang, and P. Liu. 2024. “GPTScore: Evaluate as You Desire”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 6556–76. https://doi.org/10.18653/v1/2024.naacl-long.365.
Harvard
Fu, J. et al. (2024) “GPTScore: Evaluate as You Desire”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6556–6576. Available at: https://doi.org/10.18653/v1/2024.naacl-long.365.
Vancouver
1. Fu J, Ng SK, Jiang Z, Liu P (2024) GPTScore: Evaluate as You Desire. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 6556–6576

BibTeX

@inproceedings{fu-etal-2024-gptscore,
    title = "{GPTS}core: Evaluate as You Desire",
    author = "Fu, Jinlan  and
      Ng, See-Kiong  and
      Jiang, Zhengbao  and
      Liu, Pengfei",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.365/",
    doi = "10.18653/v1/2024.naacl-long.365",
    pages = "6556--6576"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/