CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
Pei KeBosi WenAndrew FengXiao LiuXuanyu LeiJiale ChengShengyuan WangAohan ZengYuxiao DongHongning Wang
Develops an effective instruction-tuning framework and critique-generation model that automatically evaluates large language model outputs and produces detailed feedback without relying on gold-standard references or costly proprietary APIs.
Evaluating the performance of large language models has become critical as these systems are deployed across diverse applications. While proprietary models such as GPT-4 are effective automated judges, relying on commercial application programming interfaces (APIs) introduces major drawbacks, including high operational costs, latency, data leakage concerns, and vendor lock-in. Existing open-source evaluation models trained by directly prompting GPT-4 often generate vague, generic feedback, particularly in reference-free settings where no ground-truth answer is provided. Consequently, these models struggle to distinguish fine-grained quality differences between text outputs, limiting their effectiveness as automated judges.
The article demonstrates an automated data construction pipeline called Eval-Instruct and develops CritiqueLLM, an open-source model designed to produce informative, high-quality evaluation critiques for natural language generation. The framework addresses both pointwise grading (scoring a single text) and pairwise comparison (determining the superior response between two candidates), across both referenced and reference-free scenarios.
To construct high-quality training data without continuous manual annotation, the authors generated text outputs across ten diverse task categories using ten representative language models. Using human-verified GPT-4 outputs as pseudo-references, they collected detailed referenced pointwise critiques. They then applied a multi-path prompting strategy that translates referenced pointwise evaluations into pairwise and reference-free critiques through knowledge injection, reference removal, and self-reflection. A cross-validation mechanism filtered out inconsistent labels across generation paths, eliminating about 7.7% of conflicting instances. The authors then fine-tuned an open-source 6-billion-parameter model on this curated dataset across multiple evaluation settings simultaneously.
The resulting model, CritiqueLLM, achieved substantial performance gains across standard instruction-following benchmarks, such as AlignBench, AUTO-J, and LLMEval. CritiqueLLM consistently outperformed all open-source baselines and ChatGPT across text-level and system-level metrics. In system-level correlation for pointwise grading, CritiqueLLM matched the performance of GPT-4, reaching correlation scores between 0.954 and 1.000. For pairwise comparisons, CritiqueLLM achieved agreement and consistency rates that surpassed ChatGPT and other local models, aligning closely with human preferences. Furthermore, using CritiqueLLM’s critiques as feedback to refine ChatGPT outputs improved text quality across most tasks, matching the gains achieved with GPT-4 feedback, whereas ChatGPT self-critiques degraded output quality.
These findings indicate that organizations can replace costly, closed-source evaluation APIs with smaller, locally deployable models without sacrificing evaluation accuracy. High-quality automated critiques not only reduce evaluation expenditures and mitigate data privacy risks, but also provide a scalable feedback mechanism to guide continuous model refinement during training and alignment cycles.
Decision-makers should consider adopting local critique models to establish cost-effective, private evaluation pipelines. For complex reasoning, mathematics, and advanced language comprehension, automated critique feedback degraded response quality, indicating that automated critiques should currently be restricted to general knowledge, writing, and role-play tasks until reasoning-specific evaluation improves. Incorporating ensemble strategies across diverse base models is also advised to mitigate self-evaluation bias. Further development should focus on improving critique generation for step-by-step reasoning and exploring how weaker language models can better follow automated feedback.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). G-Eval establishes the rubric-guided, reference-free LLM evaluation approach that CritiqueLLM develops into detailed generated critiques.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). MT-Bench and Chatbot Arena establish key LLM-as-judge benchmarks and evaluation practices used to assess the kind of judgments CritiqueLLM targets.
- Paper: RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs, Afra Feyza Akyürek et al. (2023). RL4F first trains a dedicated model to produce natural-language feedback that improves another model’s outputs, a direct precursor to CritiqueLLM’s critique-based refinement.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). This study provides an early empirical test of LLM judgments against human evaluations, motivating the reliability and scalability problem CritiqueLLM addresses.
- Paper: Self-Generated Critiques Boost Reward Modeling for Language Models, Yue Yu 0009 et al. (2025). Critic-RM extends critique generation from evaluating outputs to training reward models that jointly produce critiques and preference scores without an external teacher.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench carries the evaluation problem into challenging reasoning, math, and coding cases, exposing limits beyond the general-generation settings CritiqueLLM handles well.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). BiGGen Bench extends LLM-based evaluation toward broad, fine-grained capability testing with instance-specific rubrics and human validation.
- Paper: JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment, Hossein A. Rahmani et al. (2025). JudgeBlender continues the push for practical open evaluation by combining smaller-model judgments to improve the robustness and cost of relevance assessment.
