Self-Generated Critiques Boost Reward Modeling for Language Models
Yue YuZhengxing ChenAston ZhangLiang TanChenguang ZhuRichard Yuanzhe PangYundi QianXuewei WangSuchin GururanganChao Zhang
Proposes Critic-RM, a framework that boosts reward modeling accuracy and data efficiency by jointly training language models to generate their own natural language critiques alongside scalar reward predictions without requiring external teacher models.
Aligning large language models with human preferences is central to deploying safe and reliable artificial intelligence. The standard reinforcement learning pipeline relies on reward models that assign a single numeric score to evaluate generated text. However, scalar scores lack interpretability, fail to utilize the natural language capabilities of modern language models, and leave systems vulnerable to reward manipulation. While language models acting as evaluators can generate detailed natural language feedback, previous methods that combine critiques with scoring require annotations from expensive, larger teacher models. Developing an effective framework that enables models to generate high-quality critiques and assign accurate scores autonomously is therefore critical.
The main objective of the article is to introduce and evaluate Critic-RM, a framework that enhances reward models by using self-generated natural language critiques without requiring supervision from external teacher models. The article evaluates whether combining self-refinement techniques with joint critique generation and reward prediction improves preference modeling accuracy, data efficiency, and reasoning error correction across diverse domains.
The researchers used an instruction-tuned 70-billion-parameter language model as the backbone to generate candidate critiques and initial quality scores for response pairs across public and synthetic datasets covering general conversation, helpfulness, reasoning, and safety. To ensure data quality without human intervention, the framework applies a two-step filtering process: it first removes candidate critiques whose ratings contradict human preference labels, and then refines the remaining critiques using model-based summarization or ranking. To address the tension between the large data volume needed for text generation and the overfitting risks of reward modeling, the authors implemented a dynamic weight schedule that trains critique generation early before shifting focus to scalar reward prediction. Evaluation was conducted across several standard and out-of-distribution benchmarks, including RewardBench and CriticBench.
The evaluation produced several key findings. First, Critic-RM outperformed standard reward models by 3.7% to 4.7% on the RewardBench benchmark and surpassed a larger 405-billion-parameter evaluation model by 6.2% to 7.3%. Second, the framework demonstrated high data efficiency; models trained on only 10% of labeled data matched or exceeded the performance of standard reward models trained on full datasets. Third, Critic-RM generalized effectively to out-of-distribution tasks, achieving an average 4% improvement over standard reward baselines and showing particular strength on complex tasks requiring multiple skills. Fourth, the generated critiques improved the reasoning correction accuracy of smaller language models by 2.5% to 3.2% compared to baseline critiques. Finally, generating multiple critiques during inference yielded further performance gains, particularly in reasoning-heavy tasks such as mathematics, coding, and safety.
These results demonstrate that self-generated critiques can substantially improve reward model accuracy and transparency without the high financial and computational costs of relying on larger teacher models. Improving reward reliability mitigates the risk of flawed model updates and provides interpretable reasoning behind automated evaluations, which is vital for high-stakes applications such as legal, clinical, or financial analysis. The findings challenge the assumption that strong teacher supervision is necessary to bootstrap critique capabilities in reward modeling.
Organizations developing or aligning language models should consider integrating self-critique generation and automated filtering pipelines into their training workflows. When deploying models under tight computational budgets, practitioners should prioritize multi-critique generation specifically for complex reasoning tasks where the performance benefits are largest. Further development should explore multi-round iterative refinement to test whether recursive self-improvement can achieve additional gains.
The study's primary limitation is that it was evaluated using a single 70-billion-parameter model family, meaning results may vary across different model architectures. Additionally, generating critiques introduces runtime overhead that increases latency during inference. Users should remain cautious of the risk that unmonitored critiques might reflect subtle underlying biases. Nevertheless, the consistent experimental gains across diverse benchmarks provide high confidence in the framework's overall effectiveness.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). Introduces the core mechanism of iterative self-critique and refinement in large language models without external supervisors, establishing the foundational principle behind Critic-RM's autonomous critique generation.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). Demonstrates the efficacy of replacing human feedback with automated AI feedback in preference learning, motivating Critic-RM's elimination of costly teacher supervision.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Establishes benchmarks and methodologies for evaluating LLMs acting as automated evaluators and judges, contextualizing the critique-and-scoring evaluation setup.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). Pioneers the use of intermediate process verification in reward modeling over simple scalar outcomes, which Critic-RM generalizes through natural language critiques.
- Paper: Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision, Collin Burns et al. (2024). Examines the theoretical dynamics of weak supervision and bootstrapping capability, providing critical insight into why self-critique works without larger teacher models.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Provides the seminal formulation for training scalar reward models to align language models with preferences, serving as the standard baseline Critic-RM seeks to enhance.
- Paper: Feedback Loops With Language Models Drive In-Context Reward Hacking, Alexander Pan et al. (2024). Details the risks of reward hacking and proxy exploitation in standard scalar feedback loops, establishing the primary vulnerability Critic-RM addresses via interpretable critiques.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Provides a comprehensive taxonomy and systematic overview of LLM-as-a-judge paradigms, extending the evaluation and critique dynamics explored in Critic-RM.
- Paper: When Can LLMs Learn to Reason with Weak Supervision?, Salman Rahman et al. (2026). Investigates how language models generalize reasoning under weak and self-supervised reward signals, complementing Critic-RM's analysis of autonomous critique generation.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). Critically analyzes self-reflection and emergent reasoning under pure reinforcement learning pipelines, building upon the self-refinement insights demonstrated in Critic-RM.
- Paper: Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability, Shobhita Sundaram et al. (2026). Extends autonomous self-improvement frameworks by using asymmetric teacher-student self-play to generate curriculum questions at the boundary of model capability.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). Theoretically unifies downstream preference learning algorithms, contextualizing how enriched reward models like Critic-RM interface with direct policy optimization.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). Explores decoupled multi-reward policy optimization, offering an algorithmic continuation for balancing multifaceted feedback signals such as critiques and scalar rewards.
