GPTScore: Evaluate as You Desire
Jinlan FuSee-Kiong NgZhengbao JiangPengfei Liu
Proposes GPTScore, a training-free evaluation framework that leverages the conditional generation probabilities of large language models to assess generated text across custom criteria defined purely through natural language instructions.
As generative artificial intelligence rapidly evolves to produce high-quality text, techniques for evaluating these outputs have failed to keep pace. Conventional automated evaluation metrics typically focus on narrow dimensions, require costly model training or human data labeling, and cannot easily adapt to customized assessment criteria required by practitioners. To overcome these constraints, the article introduces GPTScore, a text evaluation framework that scores generated content without requiring specialized model training by leveraging the conditional generation probabilities of pre-trained language models guided by natural language instructions.
The article's primary objective is to demonstrate that pre-trained language models can achieve customizable, multi-faceted, and training-free text evaluation across diverse applications. To assess this, the authors evaluated 19 pre-trained language models ranging in size from 80 million to 175 billion parameters across four major natural language generation tasks: text summarization, data-to-text generation, dialogue response generation, and machine translation. The evaluation benchmark encompassed 37 datasets across 22 distinct quality aspects, measuring automated scores against human judgments using statistical correlation benchmarks.
The analysis established several key findings. First, instructing models with explicit task and quality aspect definitions significantly improved evaluation accuracy compared to uninstructed baselines, with few-shot demonstration examples yielding further gains across most tasks. Second, training-free GPTScore models routinely matched or outperformed established supervised evaluators that underwent task-specific fine-tuning. Third, in dialogue response evaluation, the older 175-billion-parameter text-davinci-001 model drastically outperformed the human-feedback-tuned text-davinci-003, achieving an average correlation improvement of 40.8 points on turn-level assessments. Finally, enriching evaluation prompts by combining correlated quality dimensions enabled medium-sized models, such as the 6.7-billion-parameter GPT-3 Curie, to match or exceed the performance of 175-billion-parameter models.
These results demonstrate that organizations can implement highly customized, multi-dimensional quality auditing for generative text without incurring the substantial computational and annotation costs of training dedicated evaluators. However, the findings also highlight that model size and recent human-feedback tuning do not guarantee superior evaluation performance, meaning practitioners must carefully validate evaluator configurations rather than assuming newer or larger models will perform best. For production deployments, smaller and properly prompted models present a cost-effective alternative to very large commercial models.
Decision-makers aiming to adopt this framework should define clear, explicit instructions for required quality criteria and test prompt compositions that integrate related dimensions before scaling. Certain limitations must be considered: newer proprietary models such as GPT-4 were not evaluated, the internal behavior of proprietary models remains opaque, and language-model evaluators carry inherent risks of unobservable bias. The evidence strongly supports the feasibility of the framework, but organizations should conduct pilot validations and safety audits prior to high-stakes deployment.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Understanding G-Eval's framework of leveraging LLMs with chain-of-thought prompting and token probabilities for reference-free NLG evaluation directly clarifies the foundational mechanics of LLM-based zero-shot scoring explored in GPTScore.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This benchmark study formalizes the 'LLM-as-a-judge' paradigm and analyzes essential evaluative behaviors and biases that GPTScore builds upon.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). This work establishes whether prompting large models can effectively substitute for costly human evaluations across diverse rating scales, providing essential empirical justification for GPTScore.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This survey details the systemic vulnerabilities and poor human alignment of legacy NLG metrics, motivating the need for training-free, multi-faceted LLM evaluators like GPTScore.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This foundational paper introduces the emergent zero-shot and in-context instruction-following abilities of scaled generative models that GPTScore harnesses to score text.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). BLEURT illustrates the historical progression toward learned, representation-based text evaluation metrics that precede zero-shot generative scoring models.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). TRUE establishes standardized protocols and meta-evaluation principles for testing automated factual consistency metrics across natural language generation tasks.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey broadens the foundational LLM-based evaluation paradigm introduced by frameworks like GPTScore into a comprehensive taxonomy of attributes, judging architectures, and bias-mitigation strategies.
- Paper: Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models, Seungone Kim et al. (2024). Prometheus 2 advances instruction-based LLM evaluation by developing open-source language models specifically fine-tuned for custom rubric scoring and pairwise ranking.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). BiGGen Bench extends LLM-as-a-judge approaches to evaluate 103 frontier models across fine-grained capabilities using human-validated, instance-specific rubrics.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). LLMBAR critically evaluates the reliability of automated LLM judges like those utilized in GPTScore against adversarial tests containing superficial polish.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). VIEScore translates instruction-guided, training-free LLM evaluation into the multimodal visual domain to assess conditional image synthesis.
- Paper: VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation, Xuan He et al. (2024). VideoScore extends the principle of multi-aspect, model-based evaluation from generated text to synthesized video sequences across fine-grained criteria.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). ARES applies LLM scoring judges to specialized multi-faceted components of retrieval-augmented generation pipelines with prediction-powered statistical guarantees.
- Paper: Improving Automatic VQA Evaluation Using Large Language Models, Oscar Mañas et al. (2024). LAVE utilizes instruction-tuned LLM scoring to overcome rigid lexical evaluation bottlenecks specifically in open-ended visual question answering.
