LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

Helia HashemiJason EisnerCorby RossetBenjamin Van DurmeChris Kedzie

article2024ACL132 citations

Proposes LLM-RUBRIC, an automated evaluation framework that queries large language models across multidimensional rubric criteria and calibrates their probability outputs with a personalized neural network to predict individual human annotator judgments with double the accuracy of standard models.

Listen

Automated evaluation of text generated by artificial intelligence is critical for monitoring conversational agents, grading writing, and conducting legal or technical reviews at scale. Relying entirely on manual human evaluation is expensive and slow, while directly prompting large language models to produce overall quality ratings often fails because these models do not reliably align with human judgments or capture how individual human judges uniquely interpret quality.

The article demonstrates and evaluates LLM-RUBRIC, an automated evaluation framework that assesses conversational texts across multidimensional criteria and calibrates these assessments to predict the ratings of individual human evaluators. The main objective is to determine whether decomposing evaluation into fine-grained rubric questions and combining model probabilities through a personalized neural calibration network can accurately predict human assessments of overall user satisfaction.

To test this approach, the researchers authored a nine-question rubric covering specific dimensions such as naturalness, conciseness, and citation quality, as well as an overall satisfaction question on a 1-to-4 scale. They prompted a language model to evaluate texts across these individual dimensions, producing probability distributions for each choice. A lightweight feed-forward neural network then combined these multidimensional distributions to predict the individual scores of human judges. Credibility of the findings rests on an experimental design involving 741 synthetic dialogue evaluations across 24 professional annotators, followed by out-of-domain validation on 223 real human-agent conversations across 13 annotators grounded in enterprise technical queries.

The evaluation revealed several key findings. First, combining multidimensional rubric evaluations through the calibration network cut the root-mean-squared error by approximately half (reducing error below 0.5 on a 1-to-4 scale) compared to directly querying the language model for overall satisfaction. Second, the calibrated rubric approach explained approximately three-quarters of the variance in human ratings, achieving Pearson correlation coefficients of 0.35 to 0.40 against human judges, whereas raw language model predictions and factual-precision baselines achieved correlations near 0.10 to 0.22. Third, language models struggled severely on individual dimensions like conciseness and redundancy when prompted directly, but the joint calibration network compensated for these weaknesses by drawing signal across all rubric questions. Fourth, personalized weighting in the calibration network provided the largest individual performance gain, confirming that accounting for annotator differences is crucial for accurate alignment.

These findings imply that direct, uncalibrated evaluation using language models provides a misleading signal for quality and satisfaction. In contrast, decomposing evaluation into structured rubric dimensions and applying personalized calibration allows organizations to automatically monitor conversational agents with greater reliability. This framework offers practical value for developing reward signals, tracking conversational performance over time, and reducing the ongoing cost of manual human annotation in enterprise deployments.

Organizations developing or deploying conversational agents should adopt multidimensional evaluation rubrics rather than single-prompt overall quality assessments. When implementing automated scoring pipelines, teams should collect a modest baseline of human evaluations (roughly 30 evaluations per judge) to calibrate the models. Further research should focus on optimizing efficiency through adaptive rubrics that ask fewer questions dynamically and testing the system across multilingual or adversarial domains.

The primary limitations include computational overhead from querying multiple rubric dimensions and potential sensitivity to shifts in domain or user populations. Confidence in the reported results is high for English technical dialogue domains, but practitioners should exercise caution and audit for fairness before deploying automated evaluators in high-stakes or low-resource settings.

Cover for LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

Abstract

This paper introduces a framework for the automated evaluation of natural language texts. A manually constructed rubric describes how to assess multiple dimensions of interest. To evaluate a text, a large language model (LLM) is prompted with each rubric question and produces a distribution over potential responses. The LLM predictions often fail to agree well with human judges—indeed, the humans do not fully agree with one another. However, the multiple LLM distributions can be combined to predict each human judge's annotations on all questions, including a summary question that assesses overall quality or relevance. LLM-RUBRIC accomplishes this by training a small feed-forward neural network that includes both judge-specific and judge-independent parameters. When evaluating dialogue systems in a human-AI information-seeking task, we find that LLM-RUBRIC with 9 questions (assessing dimensions such as naturalness, conciseness, and citation quality) predicts human judges' assessment of overall user satisfaction, on a scale of 1–4, with RMS error < 0.5, a 2× improvement over the uncalibrated baseline.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 The LLM-RUBRIC Method
  • 3 Data
  • 3.1 Mining Topics for RAG
  • 3.2 Synthetic Dialogue Generation
  • 3.3 Real Dialogue Collection and Evaluation
  • 4 Experiments
  • 5 Results
  • 6 Analysis
  • 7 Related Work
  • 8 Conclusions
  • Acknowledgments
  • Limitations
  • Ethics Statement
  • References
  • A Aggregating Predicted Scores
  • B Handling Other Types of Datasets
  • C LLM-RUBRIC Questions
  • D Evaluation Prompt for LLM
  • E Evaluation Prompt and Preliminary Data Quality Questions for Humans
  • F Synthetic Dialogue Generation
  • G Quality of the Generated Synthetic Dialogues
  • H The User Interface for Human-Agent Dialogue Collection and Evaluation
  • I Evaluating the Collected Human-Agent Dialogues
  • J How much human judge data is needed to train calibration?
  • K Calibration Plots (Reliability Diagrams)

Knowls

  1. Knowl 1 — LLM-RUBRIC combines multidimensional LLM evidence with judge-specific calibration

    model/method

    LLM-RUBRIC automatically predicts how a particular human judge would assess a text. A manually written rubric supplies multiple-choice questions QiQ_i about different text properties, optionally including an overall-quality question. For a text TT, a fixed large language model (LLM) is prompted separately with each question and supplies probabilities for its allowed responses. A small calibration network combines the response probabilities from all questions and predicts a response distribution for each question and judge. The LLM and prompts are shared across texts and judges; the calibration network learns both shared patterns and judge-specific preferences from human annotations. The LLM probabilities over allowed responses are passed to the network without renormalizing them, so the network can in principle use probability mass assigned outside the allowed response set.

  2. Knowl 2 — Shared and judge-specific layers produce a distribution for every rubric question

    equation

    For a text TT, let QiQ_i be rubric question ii, YiY_i its finite set of allowed responses, and pLLM(yi∣T,Qi)p_{\mathrm{LLM}}(y_i\mid T,Q_i) the LLM probability assigned to response yi∈Yiy_i\in Y_i. The calibration network input xx is the concatenation of these probabilities across all rubric questions and responses. For judge aa, it computes a shared text representation and then a response distribution for each question:

    z1=σ ⁣((W1+W1a)[1;x]),z2=σ ⁣((W2+W2a)[1;z1]),z_1=\sigma\!\left((W_1+W_1^a)[1;x]\right),\qquad z_2=\sigma\!\left((W_2+W_2^a)[1;z_1]\right), p^a(yi∣T,Qi)=softmax⁡ ⁣((Vi+Via)[1;z2])yi.\hat p_a(y_i\mid T,Q_i)=\operatorname{softmax}\!\left((V_i+V_i^a)[1;z_2]\right)_{y_i}.

    Here [1;v][1;v] means the vector vv prefixed by an intercept feature; z1,z2z_1,z_2 are hidden representations; σ\sigma is the network's nonlinear activation; W1,W2,ViW_1,W_2,V_i are shared parameters; and W1a,W2a,ViaW_1^a,W_2^a,V_i^a are parameters specific to judge aa. The output for each QiQ_i is a probability distribution over YiY_i. The shared representation z2z_2 lets the prediction for one question use LLM evidence from the other rubric questions as well.

  3. Knowl 3 — Likelihood training, auxiliary-question pretraining, and mean-score decoding

    model/method

    Let D\mathcal D be the collection of human annotation records (T,i,a,yia)(T,i,a,y_i^a), where judge aa gave response yiay_i^a to question QiQ_i about text TT. LLM-RUBRIC fits the calibration network by maximizing the log-likelihood of the observed judge responses,

    ∑(T,i,a,yia)∈Dlog⁡p^a(yia∣T,Qi),\sum_{(T,i,a,y_i^a)\in\mathcal D}\log \hat p_a(y_i^a\mid T,Q_i),

    using early stopping to limit overfitting. Training has two phases: pretraining optimizes this objective over all rubric questions, then fine-tuning optimizes it using only the overall-quality question Q0Q_0. Thus, the other questions act as auxiliary tasks for learning useful representations before fitting the main prediction task. The objective treats responses to different questions as conditionally independent given the text and model.

    For a numeric response scale and squared-error evaluation, the point prediction is the distribution mean: y^ia=∑yi∈Yip^a(yi∣T,Qi)yi\hat y_i^a=\sum_{y_i\in Y_i}\hat p_a(y_i\mid T,Q_i)y_i. This minimizes expected squared error under the predicted distribution and retains a distribution that can also express uncertainty.

  4. Knowl 4 — The dialogue rubric separates satisfaction from eight assessment dimensions

    definition

    In the dialogue experiments, Q0Q_0 asks for the judge's overall satisfaction with the interaction, on a 1–4 scale. The eight auxiliary questions assess: Q1Q_1, naturalness and tone; Q2Q_2, whether the user's questions can be answered from provided references; Q3Q_3, how consistently assistant claims have citations; Q4Q_4, whether citations support the claims; Q5Q_5, whether the cited sources are the best available sources; Q6Q_6, freedom from redundancy; Q7Q_7, conciseness; and Q8Q_8, whether the number of conversational turns is appropriate for the user's information need. Responses are generally on 1–4 scales, while Q8Q_8 has three response options. Citation-related questions are inapplicable when no references are provided. The paper notes that Q8Q_8's response choices are not truly ordinal, although it nevertheless uses squared-error decoding and RMSE for that question.

  5. Knowl 5 — Synthetic Azure dialogues train the evaluator, which is tested on held-out and live interactions

    experimental setup

    The evaluation concerns English information-seeking dialogues about Microsoft Azure help topics. Topics were mined from commercial search queries using satisfactory clicks on the Azure subreddit as a filter: the resulting set contained 2,275 queries. The authors crawled 37,982 clicked URLs and, after filtering unavailable or restricted pages and applying a popularity criterion, retained 23,243 webpages; their mean length was 1,246 words with a standard deviation of 1,651.

    Using gpt-3.5-turbo-16k, five dialogue-generation setups produced 250 synthetic dialogues: 50 topics were each used with five systems ranging from no document retrieval to oracle or BM25 retrieval, with different access to topic and dialogue-history information. Three of 24 professional annotators were assigned to each dialogue; quality screening left 741 synthetic judge–dialogue annotation records. The test set of 223 real dialogues was collected from human users interacting with three live systems (no retrieval, oracle retrieval, and BM25 retrieval); each user also judged their own dialogue, and 13 annotators participated. The evaluator was trained on synthetic data, with dialogue-level five-fold cross-validation for synthetic testing, and was tested on real dialogues after training on all synthetic data. The learning-curve plot on page 26 reports reasonable performance using 20% of the synthetic training data and little further change by 80%; at full size, each judge had annotated about 30 dialogues on average.

  6. Knowl 6 — Multidimensional calibration substantially improves overall-satisfaction prediction

    data/table

    The comparison below reports prediction of human responses to overall-satisfaction question Q0Q_0. Synthetic results use five-fold cross-validation on synthetic dialogues; real-dialogue results use training on all synthetic data and testing on the 223 human-agent dialogues. RMSE is lower-is-better, while Pearson's ρ\rho, Spearman's ρ\rho, and Kendall's τ\tau measure association with human scores and are higher-is-better. The comparison table on page 7 shows that LLM-RUBRIC has the best scores among the non-oracle methods on both test sets. Against the expected raw LLM response, its RMSE falls from 0.856 to 0.396 on synthetic data and from 0.901 to 0.422 on real data; its correlations also rise substantially. The paper reports statistically significant improvements over the listed non-oracle baselines under paired permutation tests (p<0.05p<0.05). A constant mean-score predictor has RMSE 0.82 on both sets, while directly using the LLM's Q0Q_0 score performs worse than that baseline.

    Synthetic Real human-agent
    Model RMSE Pearson ρ\rho Spearman ρ\rho Kendall τ\tau RMSE Pearson ρ\rho Spearman ρ\rho Kendall τ\tau
    Random 1.499 0.002 -0.003 -0.003 1.427 0.011 0.006 0.005
    Argmax LLM Q0Q_0 0.984 0.153 0.161 0.147 1.186 0.106 0.123 0.120
    Expected LLM Q0Q_0 0.856 0.182 0.217 0.168 0.901 0.143 0.141 0.138
    Calibrated LLM Q0Q_0 0.801 0.198 0.196 0.193 0.784 0.211 0.218 0.192
    FActScore – 0.204 0.211 0.200 – 0.216 0.218 0.207
    LLM-RUBRIC 0.396 0.401 0.398 0.393 0.422 0.350 0.347 0.331
  7. Knowl 7 — Personalization and both training phases contribute to overall-score accuracy

    empirical result

    Ablations on real dialogues, with training on synthetic data, show that removing personalization causes the largest architectural degradation: RMSE/Pearson correlation changes from 0.422/0.350 for full LLM-RUBRIC to 0.601/0.198 without judge-specific parameters. Omitting pretraining gives 0.525/0.226, and omitting fine-tuning gives 0.493/0.249. Removing the LLM's Q0Q_0 probabilities gives 0.554/0.287, so the auxiliary rubric evidence is useful but does not make the direct overall-score evidence irrelevant.

    The dimension ablations show statistically significant drops (p<0.05p<0.05) when any question other than redundancy question Q6Q_6 is removed. The largest visible impact is removing citation-presence question Q3Q_3, which yields RMSE 0.573 and Pearson correlation 0.075. Removing Q8Q_8 yields 0.510/0.161, while removing Q6Q_6 yields 0.424/0.348 and is not a significant drop. These real-dialogue ablation results are reported on page 8.

  8. Knowl 8 — Calibration improves prediction of individual rubric answers, including dimensions the raw LLM misses

    data/table

    For each rubric question, the authors fine-tuned LLM-RUBRIC for that question and compared it with the expected raw LLM response on real dialogues. The table gives RMSE and Pearson correlation for each approach. LLM-RUBRIC improves both metrics on all nine questions, with reported differences statistically significant at 95% confidence (p<0.05p<0.05). Particularly large gains occur for redundancy (Q6Q_6), conciseness (Q7Q_7), and dialogue efficiency (Q8Q_8), for which the raw LLM has near-zero correlations. The page 8 results therefore show that the learned mapping can predict individual judge responses on auxiliary dimensions, not only overall satisfaction.

    Expected raw LLM LLM-RUBRIC
    Question RMSE Pearson ρ\rho RMSE Pearson ρ\rho
    Q0Q_0 0.901 0.143 0.422 0.350
    Q1Q_1 1.033 0.177 0.637 0.318
    Q2Q_2 0.799 0.140 0.543 0.265
    Q3Q_3 0.796 0.347 0.532 0.511
    Q4Q_4 0.919 0.166 0.706 0.494
    Q5Q_5 1.104 0.191 0.786 0.387
    Q6Q_6 1.726 0.030 0.430 0.279
    Q7Q_7 1.240 0.057 0.693 0.318
    Q8Q_8 0.981 0.059 0.232 0.249
  9. Knowl 9 — Overall-satisfaction probabilities are well calibrated on held-out synthetic dialogues

    empirical result

    On held-out synthetic dialogues evaluated by five-fold cross-validation, the predicted probabilities for each possible Q0Q_0 score have smoothed expected calibration error (smECE) below 0.05. The calibration plots on page 27 report smECE values of 0.014±0.0080.014\pm0.008, 0.034±0.0160.034\pm0.016, 0.044±0.0230.044\pm0.023, and 0.033±0.0220.033\pm0.022 for scores 1, 2, 3, and 4, respectively; the reported uncertainties are 95% bootstrap confidence intervals. Thus, among examples assigned a given probability to a score, its observed frequency is close to that probability on this evaluation set.

  10. Knowl 10 — Evidence for robustness is limited to synthetic-to-real transfer within the Azure task

    limitation

    The experiments establish transfer from synthetic to real dialogues only within the same English Azure information-seeking topic domain. They do not test whether a trained LLM-RUBRIC generalizes to other domains, languages, user populations, or time-varying distributions. The paper therefore treats broad-domain robustness as unverified; performance on the Azure test set is not evidence that the relationship between LLM scores and human judgments will persist under those shifts.

Coverage note — The paper's proposed extensions for adaptive question selection, dashboards, irregular or heterogeneous response types, and using predicted scores as generation rewards are omitted because they are sketches rather than experimentally evaluated contributions; detailed prompt templates and ethics discussion are also omitted because they do not add to the validated method or results.

References

  1. 1.Axel Abels, Diederik Roijers, Tom Lenaerts, Ann Nowé, and Denis Steckelmacher. 2019. Dynamic weights in multi-objective deep reinforcement learning. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 11–20.
  2. 2.Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024. MEGAVERSE: Benchmarking large language models across languages, modalities, models and tasks. In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL).
  3. 3.Mohammed Ali Al-Garadi, Sangmi Kim, Yuting Guo, Elise Warren, Yuan-Chi Yang, Sahithi Lakamana, and Abeed Sarker. 2022. Natural language model for automatic identification of intimate partner violence reports from Twitter. Array, 15.
  4. 4.Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24.
  5. 5.Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022. Stop measuring calibration when humans disagree. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1892–1915, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  6. 6.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional AI: Harmlessness from AI feedback. Computing Research Repository, arXiv:2212.08073.
  7. 7.Dwight Barry. 2017. Do not use averages with Likert scale data. Online monograph.
  8. 8.Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021. We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, pages 15–21, Online. Association for Computational Linguistics.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  10. 10.Jarosław Błasiok and Preetum Nakkiran. 2023. Smooth ece: Principled reliability diagrams via kernel smoothing. Computing Research Repository (CoRR), arXiv:2309.12236.
  11. 11.Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. 2005. The AMI meeting corpus: A pre-announcement. In Proceedings of the Second International Conference on Machine Learning for Multimodal Interaction, MLMI’05, page 28–39, Berlin, Heidelberg. Springer-Verlag.
  12. 12.Stevie Chancellor and Munmun De Choudhury. 2020. Methods in predictive techniques for mental health status on social media: A critical review. NPJ Digital Medicine, 3(1).
  13. 13.Laurent Charlin and Richard S. Zemel. 2013. The Toronto Paper Matching System: An automated paper-reviewer assignment system. In Proceedings of the ICML Workshop on Peer Reviewing and Publishing Models (PEER).
  14. 14.Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15607–15631, Toronto, Canada. Association for Computational Linguistics.
  15. 15.Eun Cheol Choi and Emilio Ferrara. 2024. FACT-GPT: Fact-checking augmentation via claim matching with LLMs. Computing Research Repository (CoRR), arXiv:2402.05904.
  16. 16.Ian Connick Covert, Wei Qiu, Mingyu Lu, Na Yoon Kim, Nathan J White, and Su-In Lee. 2023. Learning to maximize mutual information for dynamic feature selection. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 6424–6447.
  17. 17.Gianluca Detommaso, Martin Bertran, Riccardo Fogliato, and Aaron Roth. 2024. Multicalibration for confidence scoring in LLMs. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of Proceedings of Machine Learning Research, pages 10624–10641.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. Computing Research Repository, arXiv:1810.04805.
  19. 19.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. GPTScore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  20. 20.Isaac R. Galatzer-Levy, Daniel McDuff, Vivek Natarajan, Alan Karthikesalingam, and Matteo Malgaroli. 2023. The capability of large language models to measure psychiatric functioning. Computing Research Repository (CoRR), arXiv:2308.01834.
  21. 21.William Gantt, Lelia Glass, and Aaron Steven White. 2022. Decomposing and recomposing event structure. Transactions of the Association for Computational Linguistics, 10:17–34.
  22. 22.William Gantt, Benjamin Kane, and Aaron Steven White. 2020. Natural language inference with mixed effects. In Proceedings of the Ninth Joint Conference on Lexical and Computational Semantics, pages 81–87, Barcelona, Spain (Online). Association for Computational Linguistics.
  23. 23.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023. TrueTeacher: Learning factual consistency evaluation with large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2053–2070, Singapore. Association for Computational Linguistics.
  24. 24.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences of the United States of America, 120.
  25. 25.Ira Globus-Harris, Declan Harrison, Michael Kearns, Aaron Roth, and Jessica Sorrell. 2023. Multicalibration as boosting for regression. In Proceedings of the 40th International Conference on Machine Learning (ICML), volume 202 of Proceedings of Machine Learning Research, pages 11459–11492.
  26. 26.Yury Gorishniy, Ivan Rubachev, and Artem Babenko. 2022. On embeddings for numerical features in tabular deep learning. In Advances in Neural Information Processing Systems, volume 35, pages 24991–25004.
  27. 27.Donna K. Harman. 1996. Overview of the Fourth Text REtrieval Conference (TREC-4). Special Publication (NIST SP) 500-236, National Institute of Standards and Technology, Gaithersburg, Maryland.
  28. 28.He He, Hal Daumé III, and Jason Eisner. 2012. Cost-sensitive dynamic feature selection. In ICML Workshop on Inferning: Interactions between Inference and Learning, Edinburgh. 6 pages.
  29. 29.Úrsula Hébert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018. Multicalibration: Calibration for the (computationally-identifiable) masses. In Proceedings of the 35th International Conference on Machine Learning (ICML), volume 80 of Proceedings of Machine Learning Research, pages 1939–1948.
  30. 30.Tom Hosking, Phil Blunsom, and Max Bartolo. 2023. Human feedback is not gold standard. ArXiv, abs/2309.16349.
  31. 31.Andy S. Huang, Kyle Hirabayashi, Laura Barna, Deep Parikh, and Louis R. Pasquale. 2024. Assessment of a Large Language Model’s Responses to Questions and Cases About Glaucoma and Retina Management. JAMA Ophthalmology, 142(4):371–375.
  32. 32.Jiepu Jiang and James Allan. 2016. Reducing click and skip errors in search result ranking. In Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, San Francisco, CA, USA, February 22-25, 2016, pages 183–192. ACM.
  33. 33.Mohammad Kachuee, Sajad Darabi, Babak Moatamed, and Majid Sarrafzadeh. 2019. Dynamic feature acquisition using denoising autoencoders. IEEE Transactions on Neural Networks and Learning Systems, 30(8):2252–2262.
  34. 34.Nitish Shirish Keskar, Bryan McCann, Lav Varshney, Caiming Xiong, and Richard Socher. 2019. CTRL: A conditional transformer language model for controllable generation. Computing Research Repository, arXiv:1909.05858.
  35. 35.Diederik P. Kingma and Max Welling. 2019. An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 12(4):307–392.
  36. 36.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  37. 37.Siheng Li, Cheng Yang, Yichun Yin, Xinyu Zhu, Zesen Cheng, Lifeng Shang, Xin Jiang, Qun Liu, and Yujiu Yang. 2023. AutoConv: Automatically generating information-seeking conversations with large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1751–1762, Toronto, Canada. Association for Computational Linguistics.
  38. 38.Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, and Jaime Teevan. 2024. Interpretable user satisfaction estimation for conversational systems with large language models. arXiv preprint arXiv:2403.12388.
  39. 39.Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2122–2132, Austin, Texas. Association for Computational Linguistics.
  40. 40.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a. G-Eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
  41. 41.Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2023b. Calibrating LLM-based evaluator. ArXiv, abs/2309.13308.
  42. 42.Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015. The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems. In Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 285–294, Prague, Czech Republic. Association for Computational Linguistics.
  43. 43.Qingyu Lu, Baopu Qiu, Liang Ding, Liping Xie, and Dacheng Tao. 2023. Error analysis prompting enables human-like translation evaluation in large language models: A case study on ChatGPT. ArXiv, abs/2303.13809.
  44. 44.Jonathan Mellon, Jack Bailey, Ralph Scott, James Breckwoldt, Marta Miori, and Phillip Schmedeman. 2024. Do AIs know what the most important issue is? using language models to code open-text social survey responses at scale. Research & Politics, 11(1).
  45. 45.Jennifer Meyer, Thorben Jansen, Ronja Schiller, Lucas W. Liebenow, Marlene Steinbach, Andrea Horbach, and Johanna Fleckenstein. 2024. Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotions. Computers and Education: Artificial Intelligence, 6:100199.
  46. 46.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076–12100, Singapore. Association for Computational Linguistics.
  47. 47.Alexandru Niculescu-Mizil and Rich Caruana. 2005. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML).
  48. 48.OpenAI. 2024. OpenAI GPT-3.5 Turbo 16K [gpt-3.5-turbo-16k-0613]. Available at: https://platform.openai.com/docs/models/gpt-3-5-turbo.
  49. 49.Arnold Overwijk, Chenyan Xiong, and Jamie Callan. 2022. ClueWeb22: 10 billion web documents with rich information. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, page 3360–3362, New York, NY, USA. Association for Computing Machinery.
  50. 50.E. B. Page. 1968. The use of the computer in analyzing student essays. International Review of Education, 14(3):253–263.
  51. 51.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  52. 52.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  53. 53.Joan Plepi, Béla Neuendorf, Lucie Flek, and Charles Welch. 2022. Unifying data perspectivism and personalization: An application to social norms. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7391–7402, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  54. 54.Mike Quartararo, Matt Poplawski, Adam Strayer, et al. 2019. Technology Assisted Review (TAR) guidelines. Technical report, Bolch Judicial Institute of Duke Law School.
  55. 55.Alexandre Ramé, Guillaume Couairon, Mustafa Shukor, Corentin Dancette, Jean-Baptiste Gaya, Laure Soulier, and Matthieu Cord. 2023. Rewarded soups: Towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. In Advances in Neural Information Processing Systems (NeurIPS).
  56. 56.Dadi Ramesh and Suresh Kumar Sanampudi. 2022. An automated essay scoring systems: A systematic literature review. Artificial Intelligence Review, 55(3):2495–2527.
  57. 57.S. T. Roweis and L. K. Saul. 2000. Nonlinear dimensionality reduction by locally linear embedding. Science, 290(5500):2323–2326.
  58. 58.Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023. ARES: An automated evaluation framework for retrieval-augmented generation systems.
  59. 59.Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. 2023. Why don’t you do it right? analysing annotators’ disagreement in subjective tasks. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2428–2441, Dubrovnik, Croatia. Association for Computational Linguistics.
  60. 60.Naomi Saphra, Eve Fleisig, Kyunghyun Cho, and Adam Lopez. 2023. First tragedy, then parse: History repeats itself in the new era of large language models. ArXiv, abs/2311.05020.
  61. 61.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. ArXiv, abs/1707.06347.
  62. 62.Burr Settles. 2012. Active Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Springer.
  63. 63.Aviv Shamsian, Aviv Navon, Neta Glazer, Kenji Kawaguchi, Gal Chechik, and Ethan Fetaya. 2023. Auxiliary learning as an asymmetric bargaining game. In Proceedings of the 40th International Conference on Machine Learning.
  64. 64.Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022. Human evaluation of conversations is an open problem: comparing the sensitivity of various methods for evaluating dialogue agents. In Proceedings of the 4th Workshop on NLP for Conversational AI, pages 77–97, Dublin, Ireland. Association for Computational Linguistics.
  65. 65.Pradyumna Tambwekar, Murtaza Dhuliawala, Lara J. Martin, Animesh Mehta, Brent Harrison, and Mark O. Riedl. 2019. Controllable neural story plot generation via reward shaping. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, pages 5982–5988.
  66. 66.J. B. Tenenbaum, V. D. Silva, and J. C. Langford. 2000. A global geometric framework for nonlinear dimensionality reduction. Science, 290(5500):2319–2323.
  67. 67.Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024. Large language models can accurately predict searcher preferences. In 2024 International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM.
  68. 68.Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio. 2021a. SemEval-2021 task 12: Learning with disagreements. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 338–347, Online. Association for Computational Linguistics.
  69. 69.Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021b. Learning from disagreement: A survey. J. Artificial Intelligence Research, 72:1385–1470.
  70. 70.Benigno Uria, Marc-Alexandre Côté, Karol Gregor, Iain Murray, and Hugo Larochelle. 2016. Neural autoregressive distribution estimation. Journal of Machine Learning Research, 17(1):7184–7220.
  71. 71.Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021. Human evaluation of automatically generated text: Current trends and best practice guidelines. Computer Speech & Language, 67:101151.
  72. 72.Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu, and Graham Neubig. 2023. Large language models enable few-shot clustering. Computing Research Repository, arXiv:2307.00524.
  73. 73.Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. 2024. Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration. In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL).
  74. 74.Nathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, and Benjamin Van Durme. 2024. Enhancing systematic decompositional natural language inference using informal logic. Computing Research Repository (CoRR), arXiv:2402.14798.
  75. 75.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-grained human feedback gives better rewards for language model training. In Advances in Neural Information Processing Systems (NeurIPS).
  76. 76.Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. 2023. Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10967–10982, Singapore. Association for Computational Linguistics.
  77. 77.Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K. Dey, and Dakuo Wang. 2024. Mental-LLM: Leveraging large language models for mental health prediction via online text data. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT), 8(1).
  78. 78.Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan. 2019. A generalized algorithm for multi-objective reinforcement learning and policy adaptation. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 14636–14647.
  79. 79.Hiyori Yoshikawa, Tomoya Iwakura, Kimi Kaneko, Hiroaki Yoshida, Yasutaka Kumano, Kazutaka Shimada, Rafal Rzepka, and Patrycja Swieczkowska. 2021. Tell me what you read: Automatic expertise-based annotator assignment for text annotation in expert domains. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP), pages 1575–1585.
  80. 80.Xiang Yue, Boshi Wang, Ziru Chen, Kai Zhang, Yu Su, and Huan Sun. 2023. Automatic evaluation of attribution by large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 4615–4635, Singapore. Association for Computational Linguistics.
  81. 81.Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023. Conversational information seeking. Foundations and Trends® in Information Retrieval, 17(3-4):244–456.
  82. 82.Yuwei Zhang, Zihan Wang, and Jingbo Shang. 2023. ClusterLLM: Large language models as a guide for text clustering. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13903–13920, Singapore.
  83. 83.Theodore Zhao, Mu Wei, J. Samuel Preston, and Hoifung Poon. 2023. Automatic calibration and error correction for large language models via Pareto optimal self-supervision. CoRR, abs/2306.16564.
  84. 84.Ruiqi Zhong, Charlie Snell, Dan Klein, and Jason Eisner. 2023. Non-programmers can label programs indirectly via active examples: A case study with text-to-SQL. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5126–5152.
  85. 85.Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022. Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 314–324, Seattle, United States. Association for Computational Linguistics.

Citation

MLA
Hashemi, H., et al. “LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13806–34, https://doi.org/10.18653/v1/2024.acl-long.745.
APA
Hashemi, H., Eisner, J., Rosset, C., Durme, B. V., & Kedzie, C. (2024). LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13806–13834. https://doi.org/10.18653/v1/2024.acl-long.745
Chicago
Hashemi, H., J. Eisner, C. Rosset, B. V. Durme, and C. Kedzie. 2024. “LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13806–34. https://doi.org/10.18653/v1/2024.acl-long.745.
Harvard
Hashemi, H. et al. (2024) “LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13806–13834. Available at: https://doi.org/10.18653/v1/2024.acl-long.745.
Vancouver
1. Hashemi H, Eisner J, Rosset C, Durme BV, Kedzie C (2024) LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13806–13834

BibTeX

@inproceedings{hashemi-etal-2024-llm,
    title = "{LLM}-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts",
    author = "Hashemi, Helia  and
      Eisner, Jason  and
      Rosset, Corby  and
      Van Durme, Benjamin  and
      Kedzie, Chris",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.745/",
    doi = "10.18653/v1/2024.acl-long.745",
    pages = "13806--13834"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/