Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Shengwei XuYuxuan LuYifan WuJason D. HartlineGrant Schoenebeck

article2026arXiv0 citations

Demonstrates that popular LLM-as-a-judge evaluation metrics can be easily gamed despite high human correlation, introducing a mutual-information design framework that produces scoring rules resistant to strategic manipulation.

Listen

Automated evaluation metrics are central to developing and monitoring artificial intelligence systems for text generation, such as summarization, peer review, and question answering. While these metrics have traditionally served as passive scoring tools judged by their correlation with human ratings, they are increasingly used as optimization targets to steer model training and automated decisions. When metrics are used as optimization objectives, language models can game them by adopting stylistic shortcuts, verbosity, or superficial phrasing that inflates measured scores without delivering genuine, informative quality. Relying on simple statistical agreement with human judgments is therefore no longer sufficient to ensure trustworthy performance.

The article demonstrates that statistical correlation with human ratings does not guarantee robustness against strategic gaming, and it presents a systematic framework for designing and validating reference-based text evaluation metrics that resist manipulation while penalizing low-effort outputs.

To establish a common benchmark, the article introduces three test principles: correlation with human ratings (statistical alignment), sensitivity to degradation (penalizing the removal or corruption of task-relevant facts), and resistance to manipulation (preventing score inflation from superficial rephrasing, tone shifts, or empty elongation). The authors evaluated metrics across seven datasets spanning peer review, summarization, and question answering, applying six distinct degradation strategies and six manipulation strategies. Alongside existing metrics—such as lexical-overlap baselines, embedding similarity, and large language model (LLM) direct judges—the article introduced a modular mutual-information design framework that organizes metrics into four choices: information measure, estimation technique, text representation granularity, and prediction mechanism.

The findings reveal that standard evaluation approaches are highly vulnerable to exploitation. While LLM-as-a-Judge configurations achieved the strongest baseline statistical correlations with human ratings, they failed catastrophically under strategic gaming, succumbing to manipulation in 18 to 26 out of 30 test scenarios and frequently awarding higher scores to shallow surface-level summaries. In contrast, mutual-information-based metrics showed substantially greater manipulation resistance. In particular, a newly uncovered metric from the design framework—which calculates total-variation mutual information using variational estimation over statement-level decompositions—achieved zero manipulation failures out of 30 tests and failed only 3 of 37 degradation tests, while remaining highly competitive in its correlation with human ratings.

These results demonstrate that direct LLM evaluation metrics expose organizations to significant operational and safety risks when used as automated decision gates or training reward signals, as models can easily hack them with vacuous or biased text. Crucially, the experiments prove that robustness stems from the information-theoretic formulation rather than the underlying language model itself, because the same language model that failed as an absolute judge succeeded when deployed inside a mutual-information scoring rule.

Organizations evaluating or training language generation systems should avoid using standalone LLM judges as unconstrained reward functions. Decision-makers should instead deploy contrastive, mutual-information-based evaluation metrics—specifically statement-level variational scoring—to penalize uninformative reports and discount generic text. Before deploying these metrics in dynamic optimization settings, engineering teams should conduct adversarial pilot tests, implement variance-reduction methods such as multi-sample prompting or larger negative reference pools, and explore richer reference sets.

The analysis relies on fixed perturbation strategies and sample-based estimators, meaning that future adaptive models might still uncover unmodeled shortcuts. Furthermore, the variance inherent in language model critics and negative-reference sampling introduces some estimation noise. Nevertheless, the findings offer strong confidence that modular, mutual-information scoring provides a substantially more robust and strategically aligned foundation for automated text evaluation than current commercial practices.

Cover for Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Abstract

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, correlation alone is no longer sufficient: agents may strategically game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings and strategically aligned if it resists perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that decomposes existing and new metrics into four choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulation. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Mutual Information and the Data-Processing Inequality
  • 3 Test Principles
  • 4 A Design Framework for MI-Based Metrics
  • 4.1 ff-Mutual Information
  • 4.2 MI Estimation
  • 4.3 Representation
  • 4.4 Prediction Mechanism
  • 4.5 Prior Work
  • 4.6 Manipulation Robustness of the ff-MI Metrics
  • 5 Experiment
  • 5.1 Results
  • 6 Related Work
  • 7 Conclusion and Discussion
  • References
  • A Design Framework Details
  • A.1 Direct Density Ratio Estimation
  • A.2 ff-Variational Estimation
  • B TV, f-variational, Statement-level, LLM-oracle Metric
  • C Experiment Details
  • C.1 Experiment Design
  • C.2 Computational Resources
  • D Detailed Degradation and Manipulation Statistics

Knowls

  1. Knowl 1 — Unified Design Framework for Mutual-Information-Based Evaluation Metrics

    model/method

    Reference-based text evaluation metrics can be formalized through a unified four-component design tuple D:=(F,E,R,Π)\mathcal{D} := (F, E, R, \Pi), which decomposes the metric pipeline into theoretical and empirical choices:

    1. Information Measure (FF): The choice of ff-divergence defining the population mutual information objective If(X;Y)=Df(PXY∥PX⊗PY)I_f(X; Y) = D_f(P_{XY} \parallel P_X \otimes P_Y) between candidate text XX and reference text YY. Standard choices include Kullback-Leibler divergence (F=KLF = \text{KL}) and Total Variation distance (F=TVF = \text{TV}).
    2. Estimation Method (EE): The mathematical formulation used to estimate If(X;Y)I_f(X; Y), categorized into:
      • Direct estimation (E=DirectE = \text{Direct}): Estimates the joint-to-product density ratio r(x,y)=pXY(x,y)pX(x)pY(y)r(x, y) = \frac{p_{XY}(x, y)}{p_X(x)p_Y(y)} directly.
      • ff-variational estimation (E=f-variationalE = f\text{-variational}): Uses the Fenchel dual formulation to estimate divergence via a critic function T∗(x,y)∈∂f(r(x,y))T^*(x, y) \in \partial f(r(x, y)), contrasting matched candidate-reference pairs against independently drawn negative references y−∼PYy^- \sim P_Y.
    3. Text Representation (RR): The granularity map R:V∗→UR: \mathcal{V}^* \to \mathcal{U} at which reports are processed: token sequences (R=TokenR = \text{Token}), atomic statement sets (R=StatementR = \text{Statement}), or complete documents (R=Full-ReportR = \text{Full-Report}).
    4. Prediction Mechanism (Π\Pi): The model query mechanism supplying predictive quantities from represented pairs: autoregressive token likelihoods (Π=Autoregression\Pi = \text{Autoregression}) or prompting-based language model predictions (Π=LLM-oracle\Pi = \text{LLM-oracle}).

    This framework unifies prior metrics as specific configurations—such as GEM as (KL,Direct,Token,Autoregression)(\text{KL}, \text{Direct}, \text{Token}, \text{Autoregression}), GPPM-Judgment as (KL,Direct,Statement,LLM-oracle)(\text{KL}, \text{Direct}, \text{Statement}, \text{LLM-oracle}), and TVD-MI as (TV,f-variational,Full-Report,LLM-oracle)(\text{TV}, f\text{-variational}, \text{Full-Report}, \text{LLM-oracle})—while opening unexplored design combinations.

  2. Knowl 2 — Statement-Level Total Variation f-Variational Text Evaluation Metric

    model/method

    The statement-level Total Variation ff-variational metric instantiates the design tuple (TV,f-variational,Statement,LLM-oracle)(\text{TV}, f\text{-variational}, \text{Statement}, \text{LLM-oracle}). It evaluates candidate response x∈V∗x \in \mathcal{V}^* by comparing its factual support for positive reference responses against negative (cross-task) reference responses at atomic claim granularity.

    1. Statement Decomposition: For any reference text y∈V∗y \in \mathcal{V}^*, an LLM decomposes yy into a set of K(y)K(y) atomic claims: Ψ(y)={ψ1(y),…,ψK(y)(y)}\Psi(y) = \{\psi_1(y), \dots, \psi_{K(y)}(y)\}

    2. Local Support Critic: For each candidate-statement pair (x,ψk(y))(x, \psi_k(y)), an LLM judge acts as a centered binary local critic: t(x,ψk(y))={+12if x supports or aligns with statement ψk(y)−12otherwiset(x, \psi_k(y)) = \begin{cases} +\frac{1}{2} & \text{if } x \text{ supports or aligns with statement } \psi_k(y) \\ -\frac{1}{2} & \text{otherwise} \end{cases}

    3. Report-Level Critic: Local judgments are aggregated via mean pooling into a report-level critic bounded in [−12,12][-\frac{1}{2}, \frac{1}{2}]: TΨ(x,y)=1K(y)∑k=1K(y)t(x,ψk(y))T_\Psi(x, y) = \frac{1}{K(y)} \sum_{k=1}^{K(y)} t(x, \psi_k(y))

    4. Sample Contrastive Scoring: Given a set of positive references Ri+R_i^+ paired with task ii and a pool of negative references Ri−R_i^- sampled from marginal distribution PYP_Y, the score is: S^CTV(xi;Ri+,Ri−)=1∣Ri+∣∑y∈Ri+TΨ(xi,y)−1∣Ri−∣∑y−∈Ri−TΨ(xi,y−)\hat{S}_{\text{CTV}}(x_i; R_i^+, R_i^-) = \frac{1}{|R_i^+|} \sum_{y \in R_i^+} T_\Psi(x_i, y) - \frac{1}{|R_i^-|} \sum_{y^- \in R_i^-} T_\Psi(x_i, y^-)

    In expectation, the score satisfies E[S^CTV(Xi;Ri+,Ri−)]≤ITV(X;Ψ(Y))≤ITV(X;Y)\mathbb{E}[\hat{S}_{\text{CTV}}(X_i; R_i^+, R_i^-)] \le I_{\text{TV}}(X; \Psi(Y)) \le I_{\text{TV}}(X; Y), penalizing generic content that supports unrelated negative references while rewarding specific alignment with task references.

  3. Knowl 3 — Test Principles for Statistical and Strategic Alignment of Evaluation Metrics

    definition

    To assess reference-based text evaluation metrics S(x,y,y−)S(x, y, y^-), three complementary test principles are defined:

    1. Statistical Alignment (Correlation with Human Preference): Under non-adversarial conditions, metric scores si=S(xi,yi,yi−)s_i = S(x_i, y_i, y_i^-) on candidate responses xix_i must correlate with human quality ratings hih_i, quantified by Spearman's rank correlation ρ=Spearman({si}i=1n,{hi}i=1n)\rho = \text{Spearman}(\{s_i\}_{i=1}^n, \{h_i\}_{i=1}^n).
    2. Strategic Alignment — Degradation Sensitivity: When candidate text xix_i undergoes a perturbation MM that deliberately degrades, deletes, or corrupts task-relevant information (producing xiMx_i^M), the metric score should strictly decrease. Evaluated across nn instances via a one-tailed paired tt-test for negative mean score change: Δˉ(M):=1n∑i=1n(siM−si)<0\bar{\Delta}^{(M)} := \frac{1}{n} \sum_{i=1}^n (s_i^M - s_i) < 0 and quantified by the Standardized Mean Difference (SMD) dM=μM−μ(σM2+σ2)/2d_M = \frac{\mu_M - \mu}{\sqrt{(\sigma_M^2 + \sigma^2)/2}}. A metric fails if Δˉ(M)≥0\bar{\Delta}^{(M)} \ge 0 (p>0.05p > 0.05).
    3. Strategic Alignment — Manipulation Robustness: When candidate text xix_i is strategically altered via manipulation MM to game metric shortcuts (such as verbosity, sentiment, or superficial formatting) without introducing new task-specific evidence, the score should not increase. Evaluated via a paired tt-test; a metric fails if it yields a statistically significant positive score inflation Δˉ(M)>0\bar{\Delta}^{(M)} > 0 (p<0.05p < 0.05).
  4. Knowl 4 — Approximate Manipulation Robustness Bound for Estimated Mutual-Information Metrics

    theoretical result

    Let candidate text X∈XX \in \mathcal{X} and reference text Y∈YY \in \mathcal{Y} be conditionally independent given task variable WW. By the data-processing inequality for ff-divergence, for any Markov chain Y−X−σ(X)Y - X - \sigma(X) induced by candidate manipulation σ:X→X\sigma: \mathcal{X} \to \mathcal{X}: If(σ(X);Y)≤If(X;Y)I_f(\sigma(X); Y) \le I_f(X; Y)

    When mutual information IfI_f is approximated by a sample scoring function S(x,y)S(x, y), exact information monotonicity may fail due to estimation error. If the estimation error of SS is uniformly bounded by ε\varepsilon over a manipulation class Σ\Sigma, such that: sup⁡σ∈Σ∣E[S(σ(X),Y)]−If(σ(X);Y)∣≤ε\sup_{\sigma \in \Sigma} \left| \mathbb{E}[S(\sigma(X), Y)] - I_f(\sigma(X); Y) \right| \le \varepsilon

    then for every candidate manipulation σ∈Σ\sigma \in \Sigma, the expected manipulated score satisfies: E[S(σ(X),Y)]≤E[S(X,Y)]+2ε\mathbb{E}[S(\sigma(X), Y)] \le \mathbb{E}[S(X, Y)] + 2\varepsilon

    Consequently, any manipulation that reduces the true population ff-mutual information by more than 2ε2\varepsilon is guaranteed to receive a strictly lower score in expectation.

  5. Knowl 5 — Sample-Level Estimators for Information Measures in the f-MI Framework

    equation

    Let r(x,y)=pXY(x,y)pX(x)pY(y)=p(y∣x)p(y)r(x, y) = \frac{p_{XY}(x, y)}{p_X(x)p_Y(y)} = \frac{p(y \mid x)}{p(y)} be the joint-to-product density ratio between candidate xx and reference yy. The sample-level scoring functions S(x,y,y−)S(x, y, y^-) across information measures F∈{KL,TV}F \in \{\text{KL}, \text{TV}\} and estimators E∈{Direct,f-variational}E \in \{\text{Direct}, f\text{-variational}\} are:

    1. KL Direct Estimation: SKL, Direct(x,y)=log⁡r(x,y)=log⁡p(y∣x)−log⁡p(y)S_{\text{KL, Direct}}(x, y) = \log r(x, y) = \log p(y \mid x) - \log p(y)

    2. KL ff-Variational Estimation (NWJ dual bound): SKL, f-var(x,y,y−)=T∗(x,y)−exp⁡(T∗(x,y−)−1)S_{\text{KL, f-var}}(x, y, y^-) = T^*(x, y) - \exp(T^*(x, y^-) - 1) where the population-optimal critic is T∗(x,y)=1+log⁡r(x,y)T^*(x, y) = 1 + \log r(x, y).

    3. Total Variation Direct Estimation: STV, Direct(x,y)=12∣1−r(x,y)−1∣orSTV, Direct(x,y−)=12∣r(x,y−)−1∣S_{\text{TV, Direct}}(x, y) = \frac{1}{2}|1 - r(x, y)^{-1}| \quad \text{or} \quad S_{\text{TV, Direct}}(x, y^-) = \frac{1}{2}|r(x, y^-) - 1|

    4. Total Variation ff-Variational Estimation: STV, f-var(x,y,y−)=T∗(x,y)−T∗(x,y−)S_{\text{TV, f-var}}(x, y, y^-) = T^*(x, y) - T^*(x, y^-) where the population-optimal critic is T∗(x,y)=12sign(r(x,y)−1)T^*(x, y) = \frac{1}{2}\text{sign}(r(x, y) - 1).

  6. Knowl 6 — Separation Between Statistical Alignment and Strategic Manipulation Robustness

    empirical result

    Empirical evaluation across peer review (Peer Grading-WH, Peer Grading-XLSK, ICLR 2026), summarization (SummEval, SPACE), and question answering (LFQA-E, MedAESQA) reveals a structural divergence between human-rating correlation and strategic manipulation resistance:

    1. LLM-as-a-Judge Failure under Manipulation: Direct LLM judges (Claude-Haiku-4.5, Claude-Sonnet-4.5, GPT-4o-mini, GPT-5-mini) attain the highest Spearman correlations with human ratings (e.g., ρ=0.400\rho = 0.400 to 0.4770.477 on SummEval, and up to 0.6310.631 on Peer Grading-XLSK). However, they are highly vulnerable to strategic manipulations, failing 18 to 26 out of 30 manipulation tests by significantly rewarding empty elongation, tone shifts, and rephrasing.
    2. Robustness of Mutual-Information Metrics: Statement-level TV ff-variational metric (TV,f-variational,Statement,LLM-oracle)(\text{TV}, f\text{-variational}, \text{Statement}, \text{LLM-oracle}) achieves 0 failures across all 30 manipulation tests and fails only 3 of 37 degradation tests, while remaining competitive on human rating correlations (ρ=0.468\rho = 0.468 on PG-XLSK, ρ=0.208\rho = 0.208 on SummEval, ρ=0.228\rho = 0.228 on MedAESQA).
    3. Source of Robustness: The robustness of MI-based metrics stems from the contrastive, information-theoretic formulation rather than base model capability. When Claude-Haiku-4.5 is used as a direct absolute judge, it fails 20 manipulation tests; when deployed as the local critic inside statement-level TV ff-variational scoring, it passes all 30 manipulation tests.
  7. Knowl 7 — Text Representation Granularities and Prediction Mechanisms in MI-Based Metrics

    model/method

    In the (F,E,R,Π)(F, E, R, \Pi) design framework, the text representation RR and prediction mechanism Π\Pi instantiate three primary operational pipelines:

    1. Token Autoregression (R=Token,Π=AutoregressionR = \text{Token}, \Pi = \text{Autoregression}): Candidate and reference texts are represented as token sequences (y1,…,yT)(y_1, \dots, y_T). An autoregressive language model computes conditional likelihood pϕ(y∣x)=∏t=1Tpϕ(yt∣y<t,x)p_\phi(y \mid x) = \prod_{t=1}^T p_\phi(y_t \mid y_{<t}, x) and marginal likelihood pϕ(y)p_\phi(y) with the candidate omitted from the prompt, yielding Pointwise Mutual Information PMI(x;y)=log⁡pϕ(y∣x)−log⁡pϕ(y)\text{PMI}(x; y) = \log p_\phi(y \mid x) - \log p_\phi(y). This configuration requires white-box logit access and benefits from style normalization.
    2. Statement LLM-Oracle (R=Statement,Π=LLM-oracleR = \text{Statement}, \Pi = \text{LLM-oracle}): Reference texts are parsed into atomic factual statements Ψ(y)={ψ1(y),…,ψK(y)}\Psi(y) = \{\psi_1(y), \dots, \psi_K(y)\}. An LLM oracle queries support between candidate xx and each atomic claim ψk(y)\psi_k(y), and an aggregation rule A:⋃K≥1RK→RA: \bigcup_{K \ge 1} \mathbb{R}^K \to \mathbb{R} combines the per-statement values into a single prediction.
    3. Full-Report LLM-Oracle (R=Full-Report,Π=LLM-oracleR = \text{Full-Report}, \Pi = \text{LLM-oracle}): Entire reports are processed holistically. An LLM oracle outputs a discrete judgment (such as a 7-point Likert scale mapped to binned density ratios) or a binary critic decision deciding whether candidate-reference pairs originate from the same task.
  8. Knowl 8 — Perturbation Strategies for Testing Metric Degradation and Manipulation

    definition

    Twelve standardized perturbation strategies are used to benchmark the strategic alignment of text evaluation metrics:

    Degradation Strategies (Information-Decreasing):

    1. Random Replacement: Replaces the candidate response with a response from a different task in the same domain.
    2. Sentence Deletion: Drops every other sentence while retaining section headers and document format.
    3. Deletion & Completion: Deletes every other sentence and uses a helper LLM to infill gaps relying solely on remaining context.
    4. Surface Report: Generates a response using only weak context (e.g., paper abstract or news lead paragraph).
    5. Ultra-Concise Compression: Compresses text under severe length caps (e.g., 10% of length), stripping specific evidence.
    6. Opinion Flip: Reverses the evaluative verdict (e.g., accept to reject) while keeping topical vocabulary.

    Manipulation Strategies (Score-Seeking / Superficial):

    1. Rephrase: LLM rewrites the candidate in an alternate style without adding facts or altering stance.
    2. Meaningless Elongation: Appends fixed, content-free filler sentences across all candidates.
    3. Opinion Shift (Positive / Negative): Preserves topical coverage while systematically shifting stance to positive or negative sentiment.
    4. Opinion Shift (Neutral / Hedged): Flattens claims into noncommittal statements (e.g., 'is effective' →\to 'may have merit').
    5. Opinion Shift (Extreme / Strong): Amplifies statements into exaggerated superlatives to test confidence bias.
  9. Knowl 9 — Human-Rating Spearman Correlation Across Text Evaluation Metrics

    data/table

    The table below reproduces Spearman rank correlations (ρ\rho) between automatic metric scores and human quality ratings across four annotated datasets: Peer Grading-WH (PG-WH), Peer Grading-XLSK (PG-XLSK), SummEval (mean across dimensions), and MedAESQA.

    Evaluation Metric Peer review Summarization QA
    (F,E)(F, E) PG-WH PG-XLSK SummEval MedAESQA
    Autoregression, Token-representation
    (KL, Direct) (GEM) 0.460 0.475 0.138 0.113
    (KL, ff-var.) -0.218 -0.355 0.059 0.051
    (TV, Direct) 0.312 0.380 0.095 -0.025
    (TV, ff-var.) -0.053 -0.036 0.085 0.129
    LLM-oracle, Report-representation
    (KL, Direct) 0.235 0.367 0.031 0.163
    (KL, ff-var.) 0.252 0.419 0.031 0.162
    (TV, Direct) -0.231 -0.350 0.017 -0.238
    (TV, ff-var.) (TVD-MI) 0.286 0.042 0.140 0.110
    LLM-oracle, Statement-representation
    (KL, Direct) (GPPM-J) 0.254 0.307 0.073 0.041
    (KL, ff-var.) 0.281 0.251 0.075 0.064
    (TV, Direct) -0.292 -0.396 -0.043 -0.129
    (TV, ff-var.) 0.201 0.468 0.208 0.228
    Non-MI baselines
    ROUGE-L -0.244 0.167 0.111 0.151
    BLEU -0.332 0.256 0.102 0.127
    BERTScore -0.133 -0.063 0.252 0.166
    LLM-as-Judge / Claude-Haiku-4.5 0.462 0.569 0.400 0.315
    LLM-as-Judge / Claude-Sonnet-4.5 0.512 0.622 0.450 0.301
    LLM-as-Judge / GPT-5-mini 0.539 0.631 0.477 0.256
    LLM-as-Judge / GPT-4o-mini 0.492 0.528 0.452 0.208

    Standard lexical (ROUGE-L, BLEU) and embedding baselines (BERTScore) exhibit negative correlations on peer review tasks. Direct LLM judges attain the highest overall correlation scores, but statement-level (TV,f-var.)(\text{TV}, f\text{-var.}) achieves the best MI correlation on SummEval (0.208) and MedAESQA (0.228), and token-level GEM achieves the best MI correlation on peer review (0.460 and 0.475).

  10. Knowl 10 — Degradation and Manipulation Failure Counts Across Text Evaluation Metrics

    data/table

    The tables below summarize test failure counts for evaluation metrics under degradation perturbations (where failing means failing to yield a significant score drop at p<0.05p < 0.05) and manipulation perturbations (where failing means yielding a significant score increase at p<0.05p < 0.05).

    Degradation Sensitivity Failures (out of 37 total tests across 7 datasets):

    Evaluation Metric PG-WH PG-XLSK ICLR26 SummEval SPACE LFQA-E MedAESQA Total Fail
    KL-Direct-Autoreg. (GEM) 0 0 1 0 0 2 2 5
    TVD-FVar-Report 3 1 3 2 0 2 4 15
    TVD-FVar-Statement 0 0 0 1 0 1 1 3
    LLM-Judge / Claude-Haiku-4.5 0 0 1 1 0 1 2 5
    LLM-Judge / Claude-Sonnet-4.5 0 0 1 1 0 1 1 4
    LLM-Judge / GPT-4o-mini 1 3 2 1 0 1 2 10
    LLM-Judge / GPT-5-mini 0 0 1 1 0 1 2 5

    Manipulation Resistance Failures (out of 30 total tests across 7 datasets):

    Evaluation Metric PG-WH PG-XLSK ICLR26 SummEval SPACE LFQA-E MedAESQA Total Fail
    KL-Direct-Autoreg. (GEM) 1 1 5 0 0 0 0 7
    TVD-FVar-Report 0 0 0 0 0 0 0 0
    TVD-FVar-Statement 0 0 0 0 0 0 0 0
    LLM-Judge / Claude-Haiku-4.5 5 4 3 1 1 3 3 20
    LLM-Judge / Claude-Sonnet-4.5 5 4 5 1 0 2 1 18
    LLM-Judge / GPT-4o-mini 6 6 6 1 1 3 3 26
    LLM-Judge / GPT-5-mini 4 6 5 0 0 2 2 19

    Direct LLM-as-a-Judge baselines fail the majority of manipulation tests (18−2618-26 failures), whereas TV ff-variational metrics (both statement-level and full-report) achieve zero manipulation failures across all settings.

Coverage note — None was omitted; all key theoretical definitions, design tuple components, estimation formulas, test principles, empirical results, and comparative conclusions are covered.

References

  1. 1.Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mane. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
  2. 2.Chittaranjan Andrade. Mean difference, standardized mean difference (smd), and their use in meta-analysis: as simple as it gets. The Journal of clinical psychiatry, 81(5):11349, 2020.
  3. 3.Stefanos Angelidis, Reinald Kim Amplayo, Yoshihiko Suhara, Xiaolan Wang, and Mirella Lapata. Extractive opinion summarization in quantized transformer spaces. Transactions of the Association for Computational Linguistics, 9:277–293, 2021.
  4. 4.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  5. 5.Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and R Devon Hjelm. Mine: mutual information neural estimation. arXiv preprint arXiv:1801.04062, 2018.
  6. 6.Noah Burrell and Grant Schoenebeck. Measurement integrity in peer prediction: A peer assessment case study. arXiv preprint arXiv:2108.05521, 2021.
  7. 7.Siyu Chen, Jibang Wu, Yifan Wu, and Zhuoran Yang. Learning to incentivize information acquisition: Proper scoring rules meet principal-agent model. In International Conference on Machine Learning, pages 5194–5218. PMLR, 2023.
  8. 8.Yiling Chen and Fang-Yi Yu. Optimal scoring rule design under partial knowledge. In International Conference on Web and Internet Economics, pages 383–400. Springer, 2024.
  9. 9.Yiling Chen, Shi Feng, Paul Kattuman, and Fang-Yi Yu. Data reliability scoring. arXiv preprint arXiv:2510.17085, 2025.
  10. 10.Thomas M Cover. Elements of information theory. John Wiley & Sons, 1999.
  11. 11.Anirban Dasgupta and Arpita Ghosh. Crowdsourced judgement elicitation with endogenous proficiency. In Proceedings of the 22nd international conference on World Wide Web, pages 319–330, 2013.
  12. 12.Alexander R Fabbri, Wojciech Kryŕci nski, Bryan McCann, Caiming Xiong, Richard Socher, and  Dragomir Radev. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409, 2021.
  13. 13.Yuchen Fan, Chen Ling, Xin Zhong, Shuo Zhang, Heng Zhou, Yuchen Zhang, Mingyu Liang, Chengxing Xie, Ermo Hua, Zhizhou He, et al. Lfqa-e: Carefully benchmarking long-form qa evaluation. In The Fourteenth International Conference on Learning Representations, 2026.
  14. 14.Shi Feng, Hanlin Zhang, Fan Nie, Sham Kakade, and Yiling Chen. Peer-predictive self-training for language model reasoning. arXiv preprint arXiv:2604.13356, 2026.
  15. 15.Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. Bias and fairness in large language models: A survey. Computational linguistics, 50(3):1097–1179, 2024.
  16. 16.Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pages 10835–10866. PMLR, 2023.
  17. 17.Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007.
  18. 18.Deepak Gupta, Davis Bartels, and Dina Demner-Fushman. a dataset of medical questions paired with automatically generated answers and evidence-supported references. Scientific Data, 12(1): 1035, 2025.
  19. 19.Jason D Hartline, Liren Shan, Yingkai Li, and Yifan Wu. Optimal scoring rules for multi-dimensional effort. arXiv preprint arXiv:2211.03302, 2022.
  20. 20.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018.
  21. 21.Yuqing Kong. Dominantly truthful multi-task peer prediction with a constant number of tasks. In Proceedings of the fourteenth annual acm-siam symposium on discrete algorithms, pages 2398–2411. SIAM, 2020.
  22. 22.Yuqing Kong. Dominantly truthful peer prediction mechanisms with a finite number of tasks. Journal of the ACM, 71(2):1–49, 2024.
  23. 23.Yuqing Kong and Grant Schoenebeck. Water from two rocks: Maximizing the mutual information. In Proceedings of the 2018 ACM Conference on Economics and Computation, pages 177–194, 2018.
  24. 24.Yuqing Kong and Grant Schoenebeck. An information theoretic framework for designing information elicitation mechanisms that reward truth-telling. ACM Transactions on Economics and Computation (TEAC), 7(1):1–33, 2019.
  25. 25.Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg. Specification gaming: the flip side of ai ingenuity. DeepMind Blog, 3:40–53, 2020.
  26. 26.Nicolas S Lambert, David M Pennock, and Yoav Shoham. Eliciting properties of probability distributions. In Proceedings of the 9th ACM Conference on Electronic Commerce, pages 129–138, 2008.
  27. 27.Yingkai Li, Jason D Hartline, Liren Shan, and Yifan Wu. Optimization of scoring rules. In Proceedings of the 23rd ACM Conference on Economics and Computation, pages 988–989, 2022.
  28. 28.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. Transactions on Machine Learning Research, 2022.
  29. 29.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  30. 30.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), pages 3214–3252, 2022.
  31. 31.Yuxuan Lu, Shengwei Xu, Yichi Zhang, Yuqing Kong, and Grant Schoenebeck. Eliciting informative text evaluations with large language models. arXiv preprint arXiv:2405.15077, 2024.
  32. 32.Yuxuan Lu, Yifan Wu, Jason Hartline, and Michael J Curry. Aligned textual scoring rules. arXiv preprint arXiv:2507.06221, 2025.
  33. 33.John McCarthy. Measures of the value of information. Proceedings of the National Academy of Sciences, 42(9):654–655, 1956.
  34. 34.Nolan Miller, Paul Resnick, and Richard Zeckhauser. Eliciting informative feedback: The peer-prediction method. Management Science, 51(9):1359–1373, 2005.
  35. 35.Eric Neyman, Georgy Noarov, and S Matthew Weinberg. Binary scoring rules that incentivize precision. In Proceedings of the 22nd ACM Conference on Economics and Computation, pages 718–733, 2021.
  36. 36.XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory, 56(11):5847–5861, 2010.
  37. 37.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  38. 38.Vishakh Padmakumar and He He. Unsupervised extractive summarization using pointwise mutual information. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2505–2512, 2021.
  39. 39.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002.
  40. 40.Maneesha Papireddygari and Bo Waggoner. Contracts with information acquisition, via scoring rules. In Proceedings of the 23rd ACM Conference on Economics and Computation, pages 703–704, 2022.
  41. 41.Zachary Robertson and Sanmi Koyejo. Let’s measure information step-by-step: Llm-based evaluation beyond vibes. arXiv preprint arXiv:2508.05469, 2025.
  42. 42.Leonard J Savage. Elicitation of personal probabilities and expectations. Journal of the American Statistical Association, 66(336):783–801, 1971.
  43. 43.Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024.
  44. 44.Yifan Wu and Jason Hartline. Elicitationgpt: Text elicitation mechanisms via language models. arXiv preprint arXiv:2406.09363, 2024.
  45. 45.Shengwei Xu, Yichi Zhang, Paul Resnick, and Grant Schoenebeck. Spot check equivalence: an interpretable metric for information elicitation mechanisms. arXiv preprint arXiv:2402.13567, 2024.
  46. 46.Shengwei Xu, Yuxuan Lu, Grant Schoenebeck, and Yuqing Kong. Benchmarking llms’ judgments with no gold standard. In The Thirteenth International Conference on Learning Representations, 2025.
  47. 47.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019.
  48. 48.Yichi Zhang, Shengwei Xu, Grant Schoenebeck, and David Pennock. Stochastically dominant peer prediction. Advances in Neural Information Processing Systems, 38:151632–151664, 2026.
  49. 49.Shuran Zheng, Xuan Qi, Rui Ray Chen, Yongchan Kwon, and James Zou. Proper dataset valuation by pointwise mutual information. arXiv preprint arXiv:2405.18253, 2024.

Citation

MLA
Xu, S., et al. “Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics”. arXiv, 2026, https://doi.org/10.48550/arxiv.2608.01423.
APA
Xu, S., Lu, Y., Wu, Y., Hartline, J., & Schoenebeck, G. (2026). Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics. arXiv. https://doi.org/10.48550/arxiv.2608.01423
Chicago
Xu, S., Y. Lu, Y. Wu, J. Hartline, and G. Schoenebeck. 2026. “Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2608.01423.
Harvard
Xu, S. et al. (2026) “Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics”. arXiv. Available at: https://doi.org/10.48550/arxiv.2608.01423.
Vancouver
1. Xu S, Lu Y, Wu Y, Hartline J, Schoenebeck G (2026) Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics. https://doi.org/10.48550/arxiv.2608.01423

BibTeX

@misc{https://doi.org/10.48550/arxiv.2608.01423,
  doi = {10.48550/ARXIV.2608.01423},
  url = {https://arxiv.org/abs/2608.01423},
  author = {Xu, Shengwei and Lu, Yuxuan and Wu, Yifan and Hartline, Jason and Schoenebeck, Grant},
  keywords = {Artificial Intelligence (cs.AI), Computer Science and Game Theory (cs.GT), Machine Learning (cs.LG), FOS: Computer and information sciences},
  title = {Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/