keyword
LLM evaluators
An LLM evaluator is a large language model configured to assess, score, or compare the quality of text generated by other artificial intelligence systems or human users. In this paradigm, the model serves as an automated judge that reviews outputs against predefined criteria, rubrics, or reference answers to measure attributes such as factual accuracy, coherence, relevance, and safety. LLM evaluators are commonly used for pairwise comparisons to rank competing model responses as well as point-wise scoring across benchmark tasks, offering a scalable and cost-effective alternative to human annotation. While they correlate moderately well with human judgments across many natural language processing tasks, their assessments can be influenced by systemic factors such as position bias, verbosity bias, and self-enhancement bias toward their own generated outputs.
1 item

