Built independently by an author, for readers. Read the story and support ChapterPal

keyword

relevance judgment

A relevance judgment is an assessment or label that indicates the extent to which a retrieved document, text passage, or system output satisfies a specific user query or information need. In information retrieval and natural language processing evaluation, these judgments serve as the ground truth required to measure how effectively systems find and present helpful information. Evaluators assign these assessments using binary distinctions, such as relevant or non-relevant, or multi-level graded scales that capture nuances in accuracy, topicality, and utility. Collections of relevance judgments form the benchmark datasets necessary to compute standard performance metrics, including precision, recall, and normalized discounted cumulative gain, enabling consistent and reproducible comparisons across search and generative AI systems.

2 items

LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval

LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information Retrieval

Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L A Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, Emine Yilmaz

OrganizationsAmazonMicrosoftUniversity College LondonUniversity of AmsterdamUniversity of PaduaUniversity of Waterloo

Why you should read this

Presents the scope and shared task for the LLM4Eval workshop at WSDM 2025, detailing how researchers use large language models to automate relevance judgments, evaluate retrieval-augmented generation pipelines, and replace or support human assessments.

Large language models (LLMs) have demonstrated increasing task-solving abilities not present in smaller models. Utilizing the capabilities and responsibilities of LLMs for automated evaluation (LLM4Eval) has recently attracted considerable attention in multiple research communities. For instance, LLM4Eval models have been studied in the context of automated judgments, natural language generation, and retrieval augmented generation systems. We believe that the information retrieval community can significantly contribute to this growing research area by designing, implementing, analyzing, and evaluating various aspects of LLMs with applications to LLM4Eval tasks. The main goal of LLM4Eval workshop is to bring together researchers from industry and academia to discuss various aspects of LLMs for evaluation in information retrieval, including automated judgments, retrieval-augmented generation pipeline evaluation, altering human evaluation, robustness, and trustworthiness of LLMs for evaluation in addition to their impact on real-world applications. We also plan to run an automated judgment challenge prior to the workshop, where participants will be asked to generate labels for a given dataset while maximising correlation with human judgments. The format of the workshop is interactive, including roundtable and keynote sessions and tends to avoid the one-sided dialogue of a mini-conference. This is the second iteration of the workshop. The first version was held in conjunction with SIGIR 2024, attracting over 50 participants.

Added

2026-09-29

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge

Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, Huan Liu

OrganizationsArizona State UniversityEmory UniversityNorthwestern UniversityUniversity of California BerkeleyUniversity of Illinois ChicagoUniversity of Maryland, Baltimore County

Why you should read this

Presents a comprehensive taxonomy and critical overview of the "LLM-as-a-judge" paradigm, detailing its definition, methods of judgment, benchmarking, challenges, and future directions.

Assessment and evaluation have long been critical challenges in artificial intelligence (AI) and natural language processing (NLP). Traditional methods, usually matching-based or small model-based, often fall short in open-ended and dynamic scenarios. Recent advancements in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm, where LLMs are leveraged to perform scoring, ranking, or selection for various machine learning evaluation scenarios. This paper presents a comprehensive survey of LLM-based judgment and assessment, offering an in-depth overview to review this evolving field. We first provide the definition from both input and output perspectives. Then we introduce a systematic taxonomy to explore LLM-as-a-judge along three dimensions: what to judge, how to judge, and how to benchmark. Finally, we also highlight key challenges and promising future directions for this emerging area. More resources on LLM-as-a-judge are on the website: this https URL and this https URL.

Added

2026-05-20

Creative Commons License