Self-Generated Critiques Boost Reward Modeling for Language Models

Yue YuZhengxing ChenAston ZhangLiang TanChenguang ZhuRichard Yuanzhe PangYundi QianXuewei WangSuchin GururanganChao Zhang

article2025NAACL80 citations

Proposes Critic-RM, a framework that boosts reward modeling accuracy and data efficiency by jointly training language models to generate their own natural language critiques alongside scalar reward predictions without requiring external teacher models.

Listen

Aligning large language models with human preferences is central to deploying safe and reliable artificial intelligence. The standard reinforcement learning pipeline relies on reward models that assign a single numeric score to evaluate generated text. However, scalar scores lack interpretability, fail to utilize the natural language capabilities of modern language models, and leave systems vulnerable to reward manipulation. While language models acting as evaluators can generate detailed natural language feedback, previous methods that combine critiques with scoring require annotations from expensive, larger teacher models. Developing an effective framework that enables models to generate high-quality critiques and assign accurate scores autonomously is therefore critical.

The main objective of the article is to introduce and evaluate Critic-RM, a framework that enhances reward models by using self-generated natural language critiques without requiring supervision from external teacher models. The article evaluates whether combining self-refinement techniques with joint critique generation and reward prediction improves preference modeling accuracy, data efficiency, and reasoning error correction across diverse domains.

The researchers used an instruction-tuned 70-billion-parameter language model as the backbone to generate candidate critiques and initial quality scores for response pairs across public and synthetic datasets covering general conversation, helpfulness, reasoning, and safety. To ensure data quality without human intervention, the framework applies a two-step filtering process: it first removes candidate critiques whose ratings contradict human preference labels, and then refines the remaining critiques using model-based summarization or ranking. To address the tension between the large data volume needed for text generation and the overfitting risks of reward modeling, the authors implemented a dynamic weight schedule that trains critique generation early before shifting focus to scalar reward prediction. Evaluation was conducted across several standard and out-of-distribution benchmarks, including RewardBench and CriticBench.

The evaluation produced several key findings. First, Critic-RM outperformed standard reward models by 3.7% to 4.7% on the RewardBench benchmark and surpassed a larger 405-billion-parameter evaluation model by 6.2% to 7.3%. Second, the framework demonstrated high data efficiency; models trained on only 10% of labeled data matched or exceeded the performance of standard reward models trained on full datasets. Third, Critic-RM generalized effectively to out-of-distribution tasks, achieving an average 4% improvement over standard reward baselines and showing particular strength on complex tasks requiring multiple skills. Fourth, the generated critiques improved the reasoning correction accuracy of smaller language models by 2.5% to 3.2% compared to baseline critiques. Finally, generating multiple critiques during inference yielded further performance gains, particularly in reasoning-heavy tasks such as mathematics, coding, and safety.

These results demonstrate that self-generated critiques can substantially improve reward model accuracy and transparency without the high financial and computational costs of relying on larger teacher models. Improving reward reliability mitigates the risk of flawed model updates and provides interpretable reasoning behind automated evaluations, which is vital for high-stakes applications such as legal, clinical, or financial analysis. The findings challenge the assumption that strong teacher supervision is necessary to bootstrap critique capabilities in reward modeling.

Organizations developing or aligning language models should consider integrating self-critique generation and automated filtering pipelines into their training workflows. When deploying models under tight computational budgets, practitioners should prioritize multi-critique generation specifically for complex reasoning tasks where the performance benefits are largest. Further development should explore multi-round iterative refinement to test whether recursive self-improvement can achieve additional gains.

The study's primary limitation is that it was evaluated using a single 70-billion-parameter model family, meaning results may vary across different model architectures. Additionally, generating critiques introduces runtime overhead that increases latency during inference. Users should remain cautious of the risk that unmonitored critiques might reflect subtle underlying biases. Nevertheless, the consistent experimental gains across diverse benchmarks provide high confidence in the framework's overall effectiveness.

arXiv: 2411.16646
Cover for Self-Generated Critiques Boost Reward Modeling for Language Models

Abstract

Reward modeling is crucial for aligning large language models (LLMs) with human preferences, especially in reinforcement learning from human feedback (RLHF). However, current reward models mainly produce scalar scores and struggle to incorporate critiques in natural language format. We hypothesize that predicting both critiques and the scalar reward would improve reward modeling ability. Motivated by this, we propose Critic-RM, a framework that improves reward models using self-generated critiques without extra supervision. Critic-RM employs a two-stage process: generating and filtering high-quality critiques, followed by joint fine-tuning on reward prediction and critique generation. Experiments across benchmarks show that Critic-RM improves reward modeling accuracy by 3.7%–7.3% compared to standard reward models and LLM judges, demonstrating strong performance and data efficiency. Additional studies further validate the effectiveness of generated critiques in rectifying flawed reasoning steps with 2.5%–3.2% gains in improving reasoning accuracy.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Preliminaries
  • 3.2 Critique-augmented RM Training
  • 3.2.1 Critique-augmented Reward Prediction
  • 3.2.2 Rationale Generation & Filtering
  • 3.2.3 Joint Learning of Critique Generation and Reward Modeling
  • 3.3 Critic-RM Inference
  • 4 Experiments
  • 4.1 Experiment Setup: Data Generation
  • 4.2 Experiment Setup: Evaluation Datasets
  • 4.3 Baselines
  • 4.4 Implemenation Details
  • 4.5 Main Experiments: RewardBench
  • 4.6 Out-of-Distribution (OOD) Evaluation
  • 4.7 Evaluation on Critiques
  • 4.8 Data Efficiency of Reward Models
  • 4.9 Ablation Studies
  • 5 Conclusion
  • Acknowledgments
  • Limitation
  • Ethics Considerations
  • References
  • A Full Derivation Step for Eq. 4
  • B Evaluation Benchmarks
  • B.1 Evaluation Benchmarks for Reward Models
  • B.2 Evaluation Benchmarks for Critic Models
  • C Prompt Format
  • D Formatting Issues for RewardBench MATH Subset
  • E Full Results for Critic-RM
  • F Case Studies

Knowls

  1. Knowl 1 — Critic-RM Framework for Critique-Augmented Reward Modeling

    model/method

    Critic-RM is a reward modeling framework that integrates self-generated natural language critiques into scalar reward prediction using an instruction-finetuned large language model (LLM) backbone MθM_\theta, without requiring external supervision from stronger teacher models.

    Let X\mathcal{X} denote the space of input prompts and Y\mathcal{Y} denote the space of candidate responses. Critic-RM employs two heads operating on top of the shared backbone MθM_\theta:

    1. Critique Generation Model gϕ=hg∘Mθg_\phi = h_g \circ M_\theta, where hgh_g is the language modeling head inherited directly from MθM_\theta.
    2. Reward Prediction Model rψ=hr∘Mθr_\psi = h_r \circ M_\theta, where hrh_r is a scalar reward modeling head.

    Instead of scoring a response y∈Yy \in \mathcal{Y} directly given prompt x∈Xx \in \mathcal{X}, Critic-RM models critiques zz as latent intermediate variables between the response and the final scalar reward. For a preference pair consisting of prompt xx, chosen response y+y^+, and rejected response y−y^-, Critic-RM:

    1. Uses gϕg_\phi to sample candidate critiques with discrete ratings for both y+y^+ and y−y^-.
    2. Filters inconsistent instances where the average rating of y+y^+ does not exceed that of y−y^-.
    3. Refines the candidate critiques using either a summarization-based or ranking-based meta-evaluation strategy.
    4. Jointly trains MθM_\theta to generate high-quality critiques and predict scalar rewards conditioned on the concatenated response and critique [y;z][y; z].
  2. Knowl 2 — Generate-Filter-Refine Pipeline for Synthetic Critiques

    algorithm

    Critic-RM produces high-quality critique annotations from an instruction-finetuned LLM MθM_\theta using a three-step generate, filter, and refine procedure:

    Input: Training set D = {(x_i, y_i^+, y_i^-)}_{i=1}^{|D|}, LLM backbone M_theta with critique generator g_phi, candidate count N = 10, refined count K = 2
    Output: Refined critique-augmented dataset D_sub = {(x_i, y_i^+, y_i^-, Z_i^+, Z_i^-)}
    Initialize D_sub = empty_set
    for each (x, y^+, y^-) in D do
        Sample N candidate pairs for chosen response: (z_hat_i^+, s_i^+)_{i=1}^N ~ g_phi(x, y^+)
        Sample N candidate pairs for rejected response: (z_hat_i^-, s_i^-)_{i=1}^N ~ g_phi(x, y^-)
        Compute mean scores: s_bar(x, y^+) = (1 / N) * sum_{i=1}^N s_i^+ and s_bar(x, y^-) = (1 / N) * sum_{i=1}^N s_i^-
        
        if s_bar(x, y^+) > s_bar(x, y^-) then
            # Instance-level consistency check passed
            # Quality-aware critique refinement step (Option A: Summarization; Option B: Ranking)
            
            # Option A: Summarization-based Refinement
            # Generate K meta-critiques by summarizing permutations of initial critiques:
            # Z_summ = (z_k)_{k=1}^K ~ g_phi(x, y, Permute({z_hat_j}_{j=1}^N))
            
            # Option B: Ranking-based Refinement
            # Score each critique m_j ~ g_phi(x, y, z_hat_j) in [1, 10] and select top-K:
            # Z_rank = Top-K({z_hat_j}_{j=1}^N) based on scores m_j
            
            Set Z^+ and Z^- using the chosen refinement variant
            Add (x, y^+, y^-, Z^+, Z^-) to D_sub
        end if
    end for
    return D_sub
  3. Knowl 3 — Joint Training Loss and Dynamic Weight Scheduling in Critic-RM

    equation

    Critic-RM optimizes the model parameters (ϕ,ψ)(\phi, \psi) over the filtered dataset Dsub\mathcal{D}_{\text{sub}} containing tuples (x,y+,y−,Z+,Z−)(x, y^+, y^-, Z^+, Z^-) using a jointly weighted objective:

    L(ϕ,ψ)=E(x,y+,y−,Z+,Z−)∼Dsub[λ(t)⋅ℓc(ϕ)+(1−λ(t))⋅ℓr(ψ)]\mathcal{L}(\phi, \psi) = \mathbb{E}_{(x, y^+, y^-, Z^+, Z^-) \sim \mathcal{D}_{\text{sub}}} \left[ \lambda(t) \cdot \ell_c(\phi) + (1 - \lambda(t)) \cdot \ell_r(\psi) \right]

    where ℓr(ψ)\ell_r(\psi) is the critique-augmented preference loss defined via the Bradley-Terry model with sigmoid function σ\sigma:

    ℓr(x,y+,y−,z+,z−)=−log⁡σ(rψ(x,[y+;z+])−rψ(x,[y−;z−]))\ell_r(x, y^+, y^-, z^+, z^-) = -\log \sigma\left( r_\psi(x, [y^+; z^+]) - r_\psi(x, [y^-; z^-]) \right)

    and ℓc(ϕ)=ℓc(Z+;x,y+)+ℓc(Z−;x,y−)\ell_c(\phi) = \ell_c(Z^+; x, y^+) + \ell_c(Z^-; x, y^-) is the forward Kullback-Leibler (KL) divergence critique generation loss approximating the empirical oracle distribution over the KK refined critiques ZZ:

    ℓc(Z;x,y)=−1K∑z∈Zlog⁡qϕ(z∣y,x)\ell_c(Z; x, y) = -\frac{1}{K} \sum_{z \in Z} \log q_\phi(z \mid y, x)

    The dynamic weight schedule λ(t)\lambda(t) balances critique generation and reward modeling across training step tt, where TT is the number of steps per epoch, KK is the total number of training epochs (set to K=2K=2), and β=0.9\beta = 0.9 is a decay parameter:

    λ(t)={1,0<t<(K−1)T1−β×t−(K−1)TT,(K−1)T<t<KT\lambda(t) = \begin{cases} 1, & 0 < t < (K-1)T \\[6pt] 1 - \beta \times \dfrac{t - (K-1)T}{T}, & (K-1)T < t < KT \end{cases}

    This schedule forces the model to focus purely on critique language modeling during the first epoch and smoothly transition toward scalar reward prediction in the final epoch to prevent reward head overfitting.

  4. Knowl 4 — Critic-RM Inference and Inference-Time Scaling

    algorithm

    At inference time, Critic-RM evaluates a given prompt-response pair (x,y)(x, y) by generating intermediate natural language critiques before predicting the continuous scalar reward. To enhance robustness, inference-time scaling averages the predicted rewards across multiple sampled critiques:

    Input: Prompt x, Response y, Critique generator q_phi, Reward model head r_psi, Sampling temperature tau = 0.95, Number of inference critique samples m
    Output: Continuous scalar reward r
    if m == 1 then
        Sample single critique z ~ q_phi(. | x, y)
        Compute reward r = r_psi(x, [y; z])
    else
        Sample m critiques Z = {z_i}_{i=1}^m ~ q_phi(. | x, y) with temperature tau
        Compute individual rewards r_i = r_psi(x, [y; z_i]) for each z_i in Z
        Compute ensemble reward r = (1 / m) * sum_{i=1}^m r_i
    end if
    return r
  5. Knowl 5 — RewardBench Benchmark Evaluation Results

    data/table

    Critic-RM was evaluated against standard reward models, LLM judges, and critique-augmented reward models on RewardBench (2,985 test triplets). All models marked with ‡\ddagger were trained using the identical preference pairs and Llama-3.1-70B-Instruct backbone.

    Models Chat Chat_Hard Reasoning Safety Overall
    LLM-as-a-judge (For Reference)
    Llama3.1-70B-Instruct 97.2 70.2 82.8 86.0 84.0
    Llama3.1-405B-Instruct 97.2 74.6 77.6 87.1 84.1
    GPT-4-0125 95.3 74.3 87.6 86.9 86.0
    GPT-4o-0806 96.1 76.1 88.1 86.6 86.7
    Gemini-1.5-pro-0514 92.3 80.6 92.0 87.9 88.2
    Self-taught Evaluator (Iter 1) 98.3 69.0 82.6 85.7 83.9
    Self-taught Evaluator (Iter 2) 97.5 75.4 81.7 89.5 86.0
    Self-taught Evaluator (Iter 3) 96.6 84.2 91.5 81.0 88.3
    w/ inference scaling, m=32m=32 96.9 84.0 91.5 82.5 88.7
    Standard Reward Models
    RM‡\text{RM}^\ddagger 98.3 74.5 88.0 83.8 86.4
    Cohere-0514 96.4 71.3 92.3 97.7 89.4
    SteerLM-RM 70B 91.3 80.3 92.8 90.6 88.8
    Nemotron-RM 340B 95.8 87.1 91.5 93.6 92.0
    Reward Models with Critiques
    SynRM‡\text{SynRM}^\ddagger (Ours) 97.6 76.8 88.5 86.3 87.3
    CLoud‡\text{CLoud}^\ddagger (Ours) 98.0 75.6 87.6 89.0 87.6
    w/ inference scaling, m=32m=32 98.0 75.2 89.3 91.5 88.5
    Critic-RM-Summ 98.0 77.0 88.9 94.5 89.6
    w/ inference scaling, m=32m=32 97.5 77.0 91.6 95.9 90.5
    Critic-RM-Rank 97.5 79.6 90.6 94.1 90.5
    w/ inference scaling, m=32m=32 97.2 80.0 91.6 95.1 91.0

    When trained on the same preference data, Critic-RM outperforms the standard Reward Model by 3.7% to 4.7% overall accuracy and outperforms the Llama-3.1-405B judge by 6.2% to 7.3%. Inference-time scaling (m=32m=32) yields additional performance gains predominantly in reasoning-heavy subsets (Reasoning and Safety).

  6. Knowl 6 — Out-of-Distribution Reward Modeling on CrossEval, QA Feedback, and SHP

    data/table

    Critic-RM was tested on out-of-distribution (OOD) reward modeling benchmarks, including CrossEval (1,181 pairs covering individual and cross-capabilities), QA Feedback (~2,000 pairs), and Stanford Human Preferences (SHP, 3,000 subsampled pairs).

    Models CrossEval Other Datasets
    English Reasoning Coding Tool C+R T+R T+C Avg. QA Feedback SHP
    LLM-as-a-Judge (For Reference)
    Llama3.1-70B-Instruct 55.4 71.4 70.1 77.4 78.2 69.5 80.7 71.8 59.2 63.3
    Llama3.1-405B-Instruct 64.4 71.9 77.5 80.2 78.2 75.6 78.9 75.2 60.7 62.9
    Reward Models
    RM (Stiennon et al.) 59.3 72.7 70.8 75.2 68.3 72.0 72.4 70.1 58.3 65.1
    CLoud (Ankner et al.) 60.3 75.2 71.7 79.0 73.2 71.1 73.4 72.0 59.2 64.8
    Critic-RM-Summ 61.3 76.2 72.4 80.7 73.2 71.6 76.9 73.0 60.4 67.9
    Critic-RM-Rank 64.0 74.3 73.3 80.7 79.3 72.0 79.3 74.7 60.2 66.2

    Here, C+R, T+R, and T+C represent cross-capability evaluations on Code+Reasoning, Tool+Reasoning, and Tool+Code, respectively. Critic-RM outperforms standard RM baselines by an average margin of ~4%, with improvements being particularly prominent on multi-capability tasks requiring complex reasoning.

  7. Knowl 7 — Critique Quality and Reasoning Error Correction on CriticBench

    data/table

    Critic-RM was evaluated on CriticBench across 3,825 questions spanning five reasoning domains (Algorithm, Code, Symbolic, Commonsense, and Math). Critique Accuracy is measured using F1 Score, and Correction Accuracy measures the ability of a downstream policy model (Llama-3.1-8B-Instruct) to fix errors given the generated critique.

    Models Critique Accuracy (F1) Correction Accuracy (%)
    Algo Code Symb CS Math Total Algo Code Symb CS Math Total
    Baselines
    Auto-J 13B — — — — — 65.29 — — — — — —
    UltraCM 13B — — — — — 61.11 — — — — — —
    CLoud∗\text{CLoud}^* 57.22 82.87 80.56 70.18 90.35 81.91 84.75 74.56 95.35 50.22 68.48 69.56
    GPT-3.5 46.15 73.13 64.49 50.22 62.01 61.11 58.16 61.85 71.83 44.11 41.95 51.24
    GPT-4 63.51 91.36 90.75 71.56 92.55 78.75 77.66 76.29 92.41 59.96 63.57 69.96
    LLM-as-a-judge (For Reference)
    Llama3.1-70B∗\text{Llama3.1-70B}^* 60.37 84.92 86.17 65.52 88.53 80.75 77.65 76.93 88.06 59.29 57.28 66.96
    Llama3.1-405B∗\text{Llama3.1-405B}^* 86.96 88.96 90.70 72.59 93.84 86.96 86.52 81.42 90.86 63.76 63.36 72.02
    Our Model
    Critic-RM-Summ∗\textbf{Critic-RM-Summ}^* 89.79 89.36 88.36 75.26 96.09 88.25 90.55 81.89 95.82 56.95 72.54 74.33
    Critic-RM-Rank∗\textbf{Critic-RM-Rank}^* 86.13 88.88 91.10 75.02 95.49 87.93 90.42 78.44 96.43 57.39 71.77 73.87

    Models marked with ∗* use Llama-3.1-8B-Instruct as the backbone for downstream answer correction. Critic-RM achieves higher overall critique accuracy (87.93%--88.25%) than GPT-4 (78.75%) and Llama-3.1-405B (86.96%), and its feedback improves downstream correction accuracy to 73.87%--74.33%.

  8. Knowl 8 — Data Efficiency and Component Ablations of Critic-RM

    empirical result

    Ablation and data-scaling experiments on Critic-RM demonstrate the following properties:

    1. Data Efficiency: When varying training data volume (10%, 30%, 50%, 100%) on RewardBench, Critic-RM trained on only 10% of preference data outperforms the standard Bradley-Terry Reward Model trained on 100% of data (~87.5% vs ~86.4%).
    2. Loss Weight Scheduling: Replacing the dynamic weight schedule λ(t)\lambda(t) with either constant weighting across rounds or reverse scheduling (optimizing reward prediction first, followed by critique generation) causes clear accuracy drops (decreasing from ~90.5% to ~88.5%--89.0%).
    3. Training Epochs/Rounds: Training for K=2K=2 rounds provides the optimal trade-off; 1-round training yields lower accuracy (~88.5%), while 3-round training plateaus without further gain.
    4. Data Filtering and Refinement: Removing instance-level critique filtering causes a substantial performance decline (especially in the Chat-Hard subset), and removing meta-judge critique refinement (summarization or ranking) degrades accuracy from ~90.5% to below 89%.
    5. Policy Steering via Best-of-NN Sampling: When selecting among N=16N=16 candidate responses generated by Llama-3.1-Instruct (8B and 70B) on MT-Bench, Critic-RM achieves higher MT-Bench scores than standard RM and CLoud baselines.
  9. Knowl 9 — Reward Model Robustness to Response Formatting on RewardBench Math Subset

    data/table

    In the RewardBench MATH subset, human-written chosen responses consistently end with formatted answers such as \boxed{...}, whereas GPT-4-generated rejected responses end with # Answer. This difference creates a superficial formatting artifact that standard reward models can exploit.

    Models Math Original Math Rewrite
    Standard RM 80.3 76.9
    Critic-RM-Summ 83.7 83.0
    Critic-RM-Rank 82.1 81.7

    When rejected responses are rewritten so that both chosen and rejected answers use identical ending formatting, the standard RM experiences a drop of 3.4% accuracy (80.3% to 76.9%). In contrast, Critic-RM-Summ drops by only 0.7% (83.7% to 83.0%) and Critic-RM-Rank drops by 0.4% (82.1% to 81.7%), indicating that critique conditioning makes reward prediction more robust to surface-level formatting cues.

  10. Knowl 10 — Limitations of the Critic-RM Framework

    limitation

    The authors identify four main limitations of Critic-RM:

    1. Base Model Critique Capability Requirement: Critic-RM assumes that the base LLM already possesses a non-trivial level of critique generation ability. Weak base models that cannot produce coherent natural language judgments cannot bootstrap high-quality critiques through the generate-then-filter pipeline.
    2. Single Model Backbone Focus: The empirical validation was performed primarily using Llama-3.1-70B-Instruct as the backbone LLM, leaving its generalizability across diverse architectures and model scales unverified.
    3. Inference Latency Overhead: Generating full natural language critiques prior to reward score computation increases computational overhead and token latency at test time, which can restrict real-time deployment.
    4. Lack of Multi-Round Iterative Training: Critic-RM operates in a single generate-filter-train cycle rather than an iterative multi-turn self-improvement loop where the reward model iteratively refines its own generated critiques and policy over multiple rounds.

Coverage note — Prompt templates (Table 5), detailed per-task sub-breakdowns (Table 7), and qualitative case study texts (Table 8) were omitted as standalone knowls because they represent raw illustrative materials and implementation prompts rather than independent core contributions.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. 2024. Nemotron-4 340b technical report. arXiv preprint arXiv:2406.11704.
  3. 3.Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations.
  4. 4.Zachary Ankner, Mansheej Paul, Brandon Cui, Jonathan D Chang, and Prithviraj Ammanabrolu. 2024. Critique-out-loud reward models. arXiv preprint arXiv:2408.11791.
  5. 5.Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. 2024. Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. arXiv preprint arXiv:2406.18403.
  6. 6.Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345.
  7. 7.Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024a. Humans or LLMs as the judge? a study on judgement bias. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8301–8327, Miami, Florida, USA. Association for Computational Linguistics.
  8. 8.Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro. 2024b. ODIN: Disentangled reward mitigates hacking in RLHF. In Forty-first International Conference on Machine Learning.
  9. 9.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  10. 10.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2024. ULTRAFEEDBACK: Boosting language models with scaled AI feedback. In Forty-first International Conference on Machine Learning.
  11. 11.Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2024. Safe RLHF: Safe reinforcement learning from human feedback. In The Twelfth International Conference on Learning Representations.
  12. 12.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  13. 13.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback. In Thirty-seventh Conference on Neural Information Processing Systems.
  14. 14.Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022. Understanding dataset difficulty with V-usable information. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 5988–6008. PMLR.
  15. 15.Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998.
  16. 16.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  17. 17.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023. Large language models are reasoning teachers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14852–14882, Toronto, Canada. Association for Computational Linguistics.
  18. 18.Hamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi. 2024. Unpacking DPO and PPO: Disentangling best practices for learning from preference feedback. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  19. 19.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations.
  20. 20.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  21. 21.Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. 2024. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787.
  22. 22.Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Ren Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. 2024. RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback. In Forty-first International Conference on Machine Learning.
  23. 23.Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, hai zhao, and Pengfei Liu. 2024a. Generative judge for evaluating alignment. In The Twelfth International Conference on Learning Representations.
  24. 24.Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason E Weston, and Mike Lewis. 2024b. Self-alignment with instruction back-translation. In The Twelfth International Conference on Learning Representations.
  25. 25.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models.
  26. 26.Zicheng Lin, Zhibin Gou, Tian Liang, Ruilin Luo, Haowei Liu, and Yujiu Yang. 2024. CriticBench: Benchmarking LLMs for critique-correct reasoning. In Findings of the Association for Computational Linguistics ACL 2024, pages 1552–1587, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  27. 27.Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Zhe Liu, Yuan Liu, Bilal Piot, Abe Ittycheriah, Aviral Kumar, and Mohammad Saleh. 2025. RRM: Robust reward model training mitigates reward hacking. In The Thirteenth International Conference on Learning Representations.
  28. 28.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback. In Thirty-seventh Conference on Neural Information Processing Systems.
  29. 29.Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fränken, Chelsea Finn, and Alon Albalak. 2024. Generative reward models. arXiv preprint arXiv:2410.12832.
  30. 30.Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024. Llm critics help catch llm bugs. arXiv preprint arXiv:2407.00215.
  31. 31.OpenAI. 2022. Introducing ChatGPT.
  32. 32.Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. 2024. Proving test set contamination in black-box language models. In The Twelfth International Conference on Learning Representations.
  33. 33.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
  34. 34.Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason E Weston. 2024. Iterative reasoning preference optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  35. 35.Alexandre Rame, Nino Vieillard, Leonard Hussenot, Robert Dadashi, Geoffrey Cideron, Olivier Bachem, and Johan Ferret. 2024. WARM: On the benefits of weight averaged reward models. In Forty-first International Conference on Machine Learning.
  36. 36.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530.
  37. 37.William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802.
  38. 38.Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shariari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, et al. 2024. Bond: Aligning llms with best-of-n distillation. arXiv preprint arXiv:2407.14622.
  39. 39.Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, and Dong Yu. 2024. The trickle-down impact of reward inconsistency on RLHF. In The Twelfth International Conference on Learning Representations.
  40. 40.Joar Max Viktor Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming. In Advances in Neural Information Processing Systems.
  41. 41.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  42. 42.Zhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou, Zhenfang Chen, David Daniel Cox, Yiming Yang, and Chuang Gan. 2024. SALMON: Self-alignment with instructable reward models. In The Twelfth International Conference on Learning Representations.
  43. 43.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  44. 44.Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. 2024a. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 10582–10592, Miami, Florida, USA. Association for Computational Linguistics.
  45. 45.Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li. 2024b. Self-taught evaluators. arXiv preprint arXiv:2408.02666.
  46. 46.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  47. 47.Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, and Yi Dong. 2024c. Helpsteer2-preference: Complementing ratings with preferences. arXiv preprint arXiv:2410.01257.
  48. 48.Zhilin Wang, Yi Dong, Olivier Delalleau, Jiaqi Zeng, Gerald Shen, Daniel Egert, Jimmy J Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev. 2024d. Helpsteer2: Open-source dataset for training top-performing reward models. arXiv preprint arXiv:2406.08673.
  49. 49.Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev. 2024e. HelpSteer: Multi-attribute helpfulness dataset for SteerLM. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3371–3384, Mexico City, Mexico. Association for Computational Linguistics.
  50. 50.Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar. 2024. Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge. arXiv preprint arXiv:2407.19594.
  51. 51.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-grained human feedback gives better rewards for language model training. In Thirty-seventh Conference on Neural Information Processing Systems.
  52. 52.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  53. 53.Jing Nathan Yan, Tianqi Liu, Justin Chiu, Jiaming Shen, Zhen Qin, Yue Yu, Charumathi Lakshmanan, Yair Kurzion, Alexander Rush, Jialu Liu, and Michael Bendersky. 2024. Predicting text preference via structured comparative reasoning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10040–10060, Bangkok, Thailand. Association for Computational Linguistics.
  54. 54.Zihuiwen Ye, Fraser Greenlee-Scott, Max Bartolo, Phil Blunsom, Jon Ander Campos, and Matthias Gallé. 2024. Improving reward models with synthetic critiques. arXiv preprint arXiv:2405.20850.
  55. 55.Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason E Weston. 2024. Self-rewarding language models. In Forty-first International Conference on Machine Learning.
  56. 56.Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024. Evaluating large language models at evaluating instruction following. In The Twelfth International Conference on Learning Representations.
  57. 57.Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, and Rishabh Agarwal. 2024. Generative verifiers: Reward modeling as next-token prediction. arXiv preprint arXiv:2408.15240.
  58. 58.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  59. 59.Ming Zhong, Aston Zhang, Xuewei Wang, Rui Hou, Wenhan Xiong, Chenguang Zhu, Zhengxing Chen, Liang Tan, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan, Sergey Edunov, Jiawei Han, and Laurens van der Maaten. 2024. Law of the weakest link: Cross capabilities of large language models. arXiv preprint arXiv:2409.19951.
  60. 60.Banghua Zhu, Michael Jordan, and Jiantao Jiao. 2024. Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF. In Forty-first International Conference on Machine Learning.

Citation

MLA
Yu, Y., et al. “Self-Generated Critiques Boost Reward Modeling for Language Models”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 11499–514, https://doi.org/10.18653/v1/2025.naacl-long.573.
APA
Yu, Y., Chen, Z., Zhang, A., Tan, L., Zhu, C., Pang, R. Y., Qian, Y., Wang, X., Gururangan, S., Zhang, C., Kambadur, M., Mahajan, D., & Hou, R. (2025). Self-Generated Critiques Boost Reward Modeling for Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11499–11514. https://doi.org/10.18653/v1/2025.naacl-long.573
Chicago
Yu, Y., Z. Chen, A. Zhang, et al. 2025. “Self-Generated Critiques Boost Reward Modeling for Language Models”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11499–514. https://doi.org/10.18653/v1/2025.naacl-long.573.
Harvard
Yu, Y. et al. (2025) “Self-Generated Critiques Boost Reward Modeling for Language Models”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11499–11514. Available at: https://doi.org/10.18653/v1/2025.naacl-long.573.
Vancouver
1. Yu Y, Chen Z, Zhang A, et al (2025) Self-Generated Critiques Boost Reward Modeling for Language Models. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 11499–11514

BibTeX

@inproceedings{yu-etal-2025-self,
    title = "Self-Generated Critiques Boost Reward Modeling for Language Models",
    author = "Yu, Yue  and
      Chen, Zhengxing  and
      Zhang, Aston  and
      Tan, Liang  and
      Zhu, Chenguang  and
      Pang, Richard Yuanzhe  and
      Qian, Yundi  and
      Wang, Xuewei  and
      Gururangan, Suchin  and
      Zhang, Chao  and
      Kambadur, Melanie  and
      Mahajan, Dhruv  and
      Hou, Rui",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.573/",
    doi = "10.18653/v1/2025.naacl-long.573",
    pages = "11499--11514",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/