keyword
relevance scores
Relevance scores are numerical values or categorical ratings assigned to retrieved documents, items, or search outputs to quantify how effectively they satisfy a specific user query or information need. In information retrieval and search evaluation, these values can be binary indicators that categorize content simply as relevant or non-relevant, or graded multi-point scales that distinguish varying levels of usefulness, such as partially, substantially, or perfectly relevant. Assigned through human judgment, algorithmic computation, or automated model-based assessment, relevance scores serve as the primary input for evaluating search effectiveness, optimizing result ranking, and calculating performance metrics such as discounted cumulative gain.
2 items

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra
Why you should read this
Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.
The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.
Added
2026-09-30

Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin, Jaana Kekäläinen
Why you should read this
Establishes the Normalized Discounted Cumulative Gain metric to accurately evaluate search results based on multiple levels of relevance and positional decay.
Modern large retrieval environments tend to overwhelm their users by their large output. Since all documents are not of equal relevance to their users, highly relevant documents should be identified and ranked first for presentation. In order to develop IR techniques in this direction, it is necessary to develop evaluation approaches and methods that credit IR methods for their ability to retrieve highly relevant documents. This can be done by extending traditional evaluation methods, that is, recall and precision based on binary relevance judgments, to graded relevance judgments. Alternatively, novel measures based on graded relevance judgments may be developed. This article proposes several novel measures that compute the cumulative gain the user obtains by examining the retrieval result up to a given ranked position. The first one accumulates the relevance scores of retrieved documents along the ranked result list. The second one is similar but applies a discount factor to the relevance scores in order to devaluate late-retrieved documents. The third one computes the relative-to-the-ideal performance of IR techniques, based on the cumulative gain they are able to yield. These novel measures are defined and discussed and their use is demonstrated in a case study using TREC data: sample system run results for 20 queries in TREC-7. As a relevance base we used novel graded relevance judgments on a four-point scale. The test results indicate that the proposed measures credit IR methods for their ability to retrieve highly relevant documents and allow testing of statistical significance of effectiveness differences. The graphs based on the measures also provide insight into the performance IR techniques and allow interpretation, for example, from the user point of view.
Added
2026-05-08
