CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Pei KeBosi WenAndrew FengXiao LiuXuanyu LeiJiale ChengShengyuan WangAohan ZengYuxiao DongHongning Wang

article2024ACL87 citations

Develops an effective instruction-tuning framework and critique-generation model that automatically evaluates large language model outputs and produces detailed feedback without relying on gold-standard references or costly proprietary APIs.

Listen

Evaluating the performance of large language models has become critical as these systems are deployed across diverse applications. While proprietary models such as GPT-4 are effective automated judges, relying on commercial application programming interfaces (APIs) introduces major drawbacks, including high operational costs, latency, data leakage concerns, and vendor lock-in. Existing open-source evaluation models trained by directly prompting GPT-4 often generate vague, generic feedback, particularly in reference-free settings where no ground-truth answer is provided. Consequently, these models struggle to distinguish fine-grained quality differences between text outputs, limiting their effectiveness as automated judges.

The article demonstrates an automated data construction pipeline called Eval-Instruct and develops CritiqueLLM, an open-source model designed to produce informative, high-quality evaluation critiques for natural language generation. The framework addresses both pointwise grading (scoring a single text) and pairwise comparison (determining the superior response between two candidates), across both referenced and reference-free scenarios.

To construct high-quality training data without continuous manual annotation, the authors generated text outputs across ten diverse task categories using ten representative language models. Using human-verified GPT-4 outputs as pseudo-references, they collected detailed referenced pointwise critiques. They then applied a multi-path prompting strategy that translates referenced pointwise evaluations into pairwise and reference-free critiques through knowledge injection, reference removal, and self-reflection. A cross-validation mechanism filtered out inconsistent labels across generation paths, eliminating about 7.7% of conflicting instances. The authors then fine-tuned an open-source 6-billion-parameter model on this curated dataset across multiple evaluation settings simultaneously.

The resulting model, CritiqueLLM, achieved substantial performance gains across standard instruction-following benchmarks, such as AlignBench, AUTO-J, and LLMEval. CritiqueLLM consistently outperformed all open-source baselines and ChatGPT across text-level and system-level metrics. In system-level correlation for pointwise grading, CritiqueLLM matched the performance of GPT-4, reaching correlation scores between 0.954 and 1.000. For pairwise comparisons, CritiqueLLM achieved agreement and consistency rates that surpassed ChatGPT and other local models, aligning closely with human preferences. Furthermore, using CritiqueLLM’s critiques as feedback to refine ChatGPT outputs improved text quality across most tasks, matching the gains achieved with GPT-4 feedback, whereas ChatGPT self-critiques degraded output quality.

These findings indicate that organizations can replace costly, closed-source evaluation APIs with smaller, locally deployable models without sacrificing evaluation accuracy. High-quality automated critiques not only reduce evaluation expenditures and mitigate data privacy risks, but also provide a scalable feedback mechanism to guide continuous model refinement during training and alignment cycles.

Decision-makers should consider adopting local critique models to establish cost-effective, private evaluation pipelines. For complex reasoning, mathematics, and advanced language comprehension, automated critique feedback degraded response quality, indicating that automated critiques should currently be restricted to general knowledge, writing, and role-play tasks until reasoning-specific evaluation improves. Incorporating ensemble strategies across diverse base models is also advised to mitigate self-evaluation bias. Further development should focus on improving critique generation for step-by-step reasoning and exploring how weaker language models can better follow automated feedback.

Ke et al (2024).pdf
Cover for CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Abstract

The critical role of evaluation in the development of Large Language Models (LLMs) has attracted considerable attention from the research community. Most existing evaluation methods rely on comparing the generated response against the golden reference or a more powerful reference LLM, which is impractical to obtain high-quality golden references and real-time access to more powerful LLMs in real-world evaluation scenarios. In this paper, we propose CritiqueLLM, a novel critique generation model aiming to produce informative critiques that can serve as both the evaluation basis and feedback to improve LLM generation quality. We first design a point-to-batch (PtB) prompting method that enhances the evaluation performance of GPT-4 by comparing multiple responses to the same instruction. Based on this method, we construct a large-scale high-quality critique dataset containing over 1 million critiques. We then train CritiqueLLM with supervised fine-tuning, taking the critique as an intermediate step before making the final evaluation. Experimental results demonstrate that CritiqueLLM significantly outperforms existing evaluation baselines. Notably, CritiqueLLM achieves 82.5% and 91.6% accuracy on in-distribution and out-of-distribution data of the Challenge set of MT-Bench, respectively, surpassing GPT-4. Moreover, CritiqueLLM provides diagnostic feedback that can help improve the generation quality of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Task Definition and Method Overview
  • 3.2 Evaluation-Oriented Instruction Data Construction (Eval-Instruct)
  • 3.2.1 Pseudo Reference Collection
  • 3.2.2 Multi-Path Prompting
  • 3.2.3 Cross Validation
  • 3.3 Supervised Fine-Tuning
  • 4 Experiment
  • 4.1 Dataset
  • 4.2 Baselines
  • 4.3 Implementation Details
  • 4.4 Main Results
  • 4.4.1 Pointwise Grading
  • 4.4.2 Pairwise Comparison
  • 4.5 Analysis on Critique Quality
  • 4.6 Analysis of Critique as Feedback
  • 4.7 Ablation Study
  • 5 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Query Augmentation and Scoring Prompts
  • B Prompt Design for Eval-Instruct
  • C Case Study on Critique Generation

Knowls

  1. Knowl 1 — Eval-Instruct builds evaluation data from checked pseudo references

    model/method

    Eval-Instruct constructs instruction-tuning examples for four evaluation settings: pointwise grading with a reference, pointwise grading without a reference, pairwise comparison with a reference, and pairwise comparison without a reference. Pointwise critiques contain a rating and supporting explanation; pairwise critiques contain a win/tie/lose label and supporting explanation. The construction begins with queries spanning ten instruction-following task categories and responses from ten representative language models. GPT-4 responses are manually checked and revised, then used as pseudo references. GPT-4 is prompted to produce referenced pointwise critiques using those pseudo references; the method uses these critiques as the information-rich starting point for constructing the other settings. The resulting training set contains 12,102 referenced pointwise examples, 12,095 reference-free pointwise examples, 6,190 referenced pairwise examples, and 5,428 reference-free pairwise examples.

  2. Knowl 2 — Multi-path prompting transfers critique information across evaluation settings

    model/method

    Eval-Instruct uses two prompt transformations: pointwise-to-pairwise prompting injects the separate pointwise critiques for two responses into a comparison prompt, while referenced-to-reference-free prompting removes direct comparisons with the reference but asks GPT-4 to retain useful details and revise the evaluation if its explanation and label or rating do not agree.

    The transformations are applied in both possible orders. In Path 1, referenced pointwise critiques are combined into a referenced pairwise critique, which is then rewritten as reference-free pairwise evaluation. In Path 2, referenced pointwise critiques are first rewritten as reference-free pointwise critiques, which are then combined into reference-free pairwise evaluation. The pointwise-to-pairwise prompts also ask GPT-4 to reflect on inconsistencies between the pairwise decision and the order implied by the two pointwise ratings. Thus, the two paths create independent reference-free pairwise labels and critiques from the same underlying evaluation inputs.

  3. Knowl 3 — Cross-validation filters inconsistent reference-free pairwise labels

    algorithm

    Because both prompting paths produce a reference-free pairwise label for each comparison, Eval-Instruct retains a comparison only when those two labels agree. For retained comparisons, it includes the critiques from both paths in the training data; comparisons with conflicting labels are discarded. This filtering removed 7.7% of the candidate reference-free pairwise data. In an ablation, removing this cross-validation step reduced pairwise agreement from 70.56% to 66.13% in the referenced AlignBench setting and from 58.81% to 57.44% in the reference-free setting.

  4. Knowl 4 — CRITIQUELLM is fine-tuned as a unified four-setting critique generator

    model/method

    CRITIQUELLM is produced by supervised fine-tuning ChatGLM3-6B on all four Eval-Instruct datasets in a multi-task setup. The model is trained to generate each setting's critique from its corresponding evaluation input: query and response, with or without a reference for pointwise grading; and query plus two responses, with or without a reference for pairwise comparison. Pairwise training examples are augmented by swapping the response order and making the corresponding changes to the critique. The reported training configuration uses 8 A800 GPUs, AdamW with weight decay 0.1, a peak learning rate of 6e-5, 10% warmup, maximum sequence length 8,192, batch size 64, and 5 epochs. Greedy decoding is used for the main results.

  5. Knowl 5 — Pointwise grading approaches GPT-4 at system level on AlignBench

    empirical result

    On AlignBench, pointwise grading is evaluated by correlations between model scores and human judgments. Text-level correlations average per-query correlations across generated responses; system-level correlations compare each generation model's average score with its human score. The metrics are Pearson r, Spearman ρ, and Kendall τ. In the order r, ρ, τ, CRITIQUELLM obtains text-level correlations of 0.555, 0.523, 0.477 with references and 0.366, 0.352, 0.319 without references. Its system-level results are 0.995, 1.000, 1.000 with references and 0.954, 0.976, 0.929 without references. GPT-4 obtains 0.629, 0.583, 0.532 and 0.523, 0.494, 0.447 at text level, and 0.995, 1.000, 1.000 and 0.997, 0.976, 0.929 at system level, for the same respective settings. CRITIQUELLM exceeds the reported open-source baselines on all these pointwise correlations and nearly matches GPT-4 at system level.

  6. Knowl 6 — Pairwise comparison performs strongly across three benchmarks

    empirical result

    Pairwise performance is measured by agreement and consistency. For each example, the evaluator is run twice with the response order swapped; agreement requires the decisions to be consistent with each other and with the human preference label, while consistency measures whether the two order-swapped decisions agree. CRITIQUELLM's agreement/consistency percentages are 70.56/89.25 for referenced AlignBench, 58.81/83.06 for reference-free AlignBench, 50.93/82.76 for reference-free AUTO-J (Eval-P), and 50.72/85.95 for reference-free LLMEval. It exceeds ChatGPT and all reported open-source baselines on both metrics in these settings. GPT-4's corresponding values are 74.69/86.75, 70.25/84.88, 62.28/86.28, and 50.98/84.71; CRITIQUELLM therefore has higher consistency than GPT-4 in all four settings, but lower agreement in each.

  7. Knowl 7 — Generated critiques are preferred to ChatGPT critiques and competitive with GPT-4

    empirical result

    The authors sampled 100 pairwise-comparison inputs from three benchmarks and compared CRITIQUELLM critiques against critiques from ChatGPT, GPT-4, and CRITIQUELLM trained on data obtained by directly prompting GPT-4 (CRITIQUELLM (DP)). GPT-4 judged each pair for correctness, helpfulness, and informativeness, in that priority order, and humans checked the judgments. The plotted win/tie/lose percentages for CRITIQUELLM are 46/29/25 against ChatGPT, 31/35/34 against GPT-4, and 37/41/22 against CRITIQUELLM (DP). These results favor CRITIQUELLM over ChatGPT and the direct-prompting-data alternative, while showing a near-even comparison with GPT-4.

  8. Knowl 8 — CRITIQUELLM critiques improve ChatGPT outputs overall, with task-specific regressions

    empirical result

    In a feedback experiment, ChatGPT generated reference-free critiques of its own responses, GPT-4 generated reference-free critiques, or CRITIQUELLM generated reference-free critiques; ChatGPT then revised its response using the critique. GPT-4 assigned referenced pointwise scores to the original and revised responses on AlignBench. The overall score was 6.385 before revision, 6.300 after using ChatGPT critiques, 6.545 after using GPT-4 critiques, and 6.530 after using CRITIQUELLM critiques. For CRITIQUELLM feedback, scores rose in open-ended QA (7.000 to 7.381), professional tasks (5.824 to 6.765), fundamental tasks (6.310 to 6.414), role play (7.260 to 7.407), and writing (7.154 to 7.192); they fell in logical reasoning (5.318 to 5.136), mathematics (6.160 to 6.000), and Chinese understanding (6.000 to 5.315). The authors report that misleading critiques of reasoning processes contributed to declines on the harder reasoning-related tasks.

  9. Knowl 9 — Ablations show the value of explanations and decoding trade-offs

    empirical result

    On AlignBench, the greedy-decoding CRITIQUELLM baseline has text-level Pearson r of 0.555/0.366 for referenced/reference-free pointwise grading and pairwise agreement of 70.56%/58.81%. Removing explanations from training critiques lowers these metrics to 0.509/0.332 and 60.19%/51.56%, respectively. Removing cross-validation yields 0.566/0.361 and 66.13%/57.44%: the pointwise effect is mixed, while pairwise agreement declines in both settings. Self-consistency decoding, using five candidate critiques, gives the highest listed pointwise correlations (0.573/0.384) but lower pairwise agreement (69.13%/58.44%) than greedy decoding. Beam search (beam size 4) yields 0.554/0.374 and 70.31%/57.75%; nucleus sampling with temperature and p both 0.9 yields 0.547/0.353 and 68.69%/57.31%. The results show that explanations help both evaluation tasks, while decoding choices trade off pointwise and pairwise performance.

  10. Knowl 10 — Prompting cost and evaluator bias remain limitations

    limitation

    The authors note that multi-path prompting lengthens the prompts and increases API costs for constructing the evaluation data. They also report that CRITIQUELLM, like other model-based evaluators, can favor responses produced by the same base model, a form of self-evaluation or self-enhancement bias. They suggest using evaluators with different base models to reduce dependence on any one evaluator, but leave a satisfactory solution to this bias for future work. The feedback experiment also exposes a limitation on difficult reasoning tasks: CRITIQUELLM can produce misleading critiques of reasoning processes, and revisions based on those critiques can reduce response scores.

Coverage note — No substantial contributed material was omitted. The detailed prompt templates and illustrative critique cases are omitted because they instantiate the methods and findings captured above rather than adding separate results.

References

  1. 1.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  3. 3.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72.
  4. 4.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv preprint arXiv:2006.14799.
  5. 5.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109.
  6. 6.Kai Chen, Chunwei Wang, Kuo Yang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang, and Qun Liu. 2024. Gaining wisdom from setbacks: Aligning large language models via mistake analysis. In The Twelfth International Conference on Learning Representations.
  7. 7.Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the use of large language models for reference-free text quality evaluation: A preliminary empirical study. arXiv preprint arXiv:2304.00723.
  8. 8.Pierre Jean A Colombo, Chloé Clavel, and Pablo Piantanida. 2022. Infolm: A new metric to evaluate summarization & data2text generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 10554–10562.
  9. 9.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback. arXiv preprint arXiv:2310.01377.
  10. 10.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  11. 11.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  12. 12.Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. Roscoe: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations.
  13. 13.Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov. 2023. On the blind spots of model-based evaluation metrics for text generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12067–12097.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations.
  15. 15.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In 8th International Conference on Learning Representations.
  16. 16.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023a. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798.
  17. 17.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023b. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322.
  18. 18.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  19. 19.Pei Ke, Fei Huang, Fei Mi, Yasheng Wang, Qun Liu, Xiaoyan Zhu, and Minlie Huang. 2023. DecompEval: Evaluating generated texts as unsupervised decomposed question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9676–9691.
  20. 20.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations.
  21. 21.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations.
  22. 22.Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Huang. 2023. A systematic study and comprehensive evaluation of ChatGPT on benchmark datasets. In Findings of the Association for Computational Linguistics: ACL 2023, pages 431–469.
  23. 23.Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2024. Generative judge for evaluating alignment. In The Twelfth International Conference on Learning Representations.
  24. 24.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  25. 25.Xiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Zhuoer Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, et al. 2023a. Alignbench: Benchmarking chinese alignment of large language models. arXiv preprint arXiv:2311.18743.
  26. 26.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2024. Agentbench: Evaluating llms as agents. In The Twelfth International Conference on Learning Representations.
  27. 27.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522.
  28. 28.OpenAI. 2022. Introducing chatgpt.
  29. 29.OpenAI. 2023. GPT-4 technical report. arXiv preprint arXiv:2303.08774.
  30. 30.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318.
  31. 31.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, page 20.
  32. 32.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3505–3506.
  33. 33.Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati. 2023. Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems. arXiv preprint arXiv:2310.12397.
  34. 34.Shichao Sun, Junlong Li, Weizhe Yuan, Ruifeng Yuan, Wenjie Li, and Pengfei Liu. 2024. The critique of critique. arXiv preprint arXiv:2401.04518.
  35. 35.Tianxiang Sun, Xiaotian Zhang, Zhengfu He, Peng Li, Qinyuan Cheng, Hang Yan, Xiangyang Liu, Yunfan Shao, Qiong Tang, Xingjian Zhao, Ke Chen, Yining Zheng, Zhejian Zhou, Ruixiao Li, Jun Zhan, Yunhua Zhou, Linyang Li, Xiaogui Yang, Lingling Wu, Zhangyue Yin, Xuanjing Huang, and Xipeng Qiu. 2023. Moss: Training conversational language models from synthetic data.
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  37. 37.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  38. 38.Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023. Can large language models really improve by self-critiquing their own plans? arXiv preprint arXiv:2310.08118.
  39. 39.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023a. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  40. 40.Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O’Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023b. Shepherd: A critic for language model generation. arXiv preprint arXiv:2308.04592.
  41. 41.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023c. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  42. 42.Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2024. Pandalm: An automatic evaluation benchmark for LLM instruction tuning optimization. In The Twelfth International Conference on Learning Representations.
  43. 43.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023d. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 13484–13508.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837.
  45. 45.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, et al. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  46. 46.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: an open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations.
  47. 47.Yue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. Llmeval: A preliminary study on how to evaluate large language models. In The 38th Annual AAAI Conference on Artificial Intelligence.
  48. 48.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023. Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.
  49. 49.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  50. 50.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  51. 51.Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631.

Citation

MLA
Ke, P., et al. “CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13034–54, https://doi.org/10.18653/v1/2024.acl-long.704.
APA
Ke, P., Wen, B., Feng, A., Liu, X., Lei, X., Cheng, J., Wang, S., Zeng, A., Dong, Y., Wang, H., Tang, J., & Huang, M. (2024). CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13034–13054. https://doi.org/10.18653/v1/2024.acl-long.704
Chicago
Ke, P., B. Wen, A. Feng, et al. 2024. “CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13034–54. https://doi.org/10.18653/v1/2024.acl-long.704.
Harvard
Ke, P. et al. (2024) “CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13034–13054. Available at: https://doi.org/10.18653/v1/2024.acl-long.704.
Vancouver
1. Ke P, Wen B, Feng A, et al (2024) CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13034–13054

BibTeX

@inproceedings{ke-etal-2024-critiquellm,
    title = "{C}ritique{LLM}: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation",
    author = "Ke, Pei  and
      Wen, Bosi  and
      Feng, Andrew  and
      Liu, Xiao  and
      Lei, Xuanyu  and
      Cheng, Jiale  and
      Wang, Shengyuan  and
      Zeng, Aohan  and
      Dong, Yuxiao  and
      Wang, Hongning  and
      Tang, Jie  and
      Huang, Minlie",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.704/",
    doi = "10.18653/v1/2024.acl-long.704",
    pages = "13034--13054"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/