Built independently by an author, for readers. Read the story and support ChapterPal

keyword

relevance judgments

Relevance judgments are assessments that determine the extent to which a document, passage, or retrieved item satisfies a specific search query or information need. In the field of information retrieval, these annotations serve as the ground truth required to train ranking algorithms and measure the retrieval quality of search systems. Traditionally compiled by human assessors through manual inspection, relevance judgments can be formatted as binary labels indicating whether an item is relevant or non-relevant, or as graded scores across multi-point scales reflecting varying degrees of relevance. Within standardized test collections, these reference labels allow researchers and engineers to reliably benchmark system performance using metrics such as precision, recall, and discounted cumulative gain.

5 items

SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas

OrganizationsAlan Turing Institute & AmazonMicrosoftUniversity College LondonUniversity of Sheffield

Why you should read this

Presents SynDL, a large-scale passage retrieval benchmark extending the TREC Deep Learning Track across more than 1,900 queries with synthetic language model labels, providing a cost-effective evaluation resource that yields search system rankings highly correlated with human assessments.

Large-scale test collections play a crucial role in Information Retrieval (IR) research. However, according to the Cranfield paradigm and the research into publicly available datasets, the existing information retrieval research studies are commonly developed on small-scale datasets that rely on human assessors for relevance judgments - a time-intensive and expensive process. Recent studies have shown the strong capability of Large Language Models (LLMs) in producing reliable relevance judgments with human accuracy but at a greatly reduced cost. In this paper, to address the missing large-scale ad-hoc document retrieval dataset, we extend the TREC Deep Learning Track (DL) test collection via additional language model synthetic labels to enable researchers to test and evaluate their search systems at a large scale. Specifically, such a test collection includes more than 1,900 test queries from the previous years of tracks. We compare system evaluation with past human labels from past years and find that our synthetically created large-scale test collection can lead to highly correlated system rankings.

Added

2026-09-30

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra

OrganizationsAmazonMicrosoftThe Alan Turing InstituteUniversity College London

Why you should read this

Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.

The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.

Added

2026-09-30

Learning to rank: from pairwise approach to listwise approach

Learning to rank: from pairwise approach to listwise approach

Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, Hang Li

OrganizationsMicrosoftNational Taiwan UniversityTsinghua University

Why you should read this

Proposes the ListNet algorithm that directly optimizes a permutation probability based cross entropy loss to demonstrate the theoretical advantages of evaluating the entire ranked list.

The paper is concerned with learning to rank, which is is to construct a model or a function for ranking objects. Learning to rank is useful for document retrieval, collaborative filtering, and many other applications. Several methods for learning to rank have been proposed, which take object pairs as ‘instances’ in learning. We refer to them as the pairwise approach in this paper. Although the pairwise approach offers advantages, it ignores the fact that ranking is a prediction task on list of objects. The paper postulates that learning to rank should adopt the listwise approach in which lists of objects are used as ‘instances’ in learning. The paper proposes a new probabilistic method for the approach. Specifically it introduces two probability models, respectively referred to as permutation probability and top one probability, to define a listwise loss function for learning. Neural Network and Gradient Descent are then employed as model and algorithm in the learning method. Experimental results on information retrieval show that the proposed listwise approach performs better than the pairwise approach.

Added

2026-05-22

License

Published with permission

Optimizing search engines using clickthrough data

Optimizing search engines using clickthrough data

Thorsten Joachims

OrganizationsCornell University

Why you should read this

Proposes a method for transforming user clickstream logs into relative preference judgments that can be directly used to train ranking support vector machines.

This paper presents an approach to automatically optimizing the retrieval quality of search engines using clickthrough data. Intuitively, a good information retrieval system should present relevant documents high in the ranking, with less relevant documents following below. While previous approaches to learning retrieval functions from examples exist, they typically require training data generated from relevance judgments by experts. This makes them difficult and expensive to apply. The goal of this paper is to develop a method that utilizes clickthrough data for training, namely the query-log of the search engine in connection with the log of links the users clicked on in the presented ranking. Such clickthrough data is available in abundance and can be recorded at very low cost. Taking a Support Vector Machine (SVM) approach, this paper presents a method for learning retrieval functions. From a theoretical perspective, this method is shown to be well-founded in a risk minimization framework. Furthermore, it is shown to be feasible even for large sets of queries and features. The theoretical results are verified in a controlled experiment. It shows that the method can effectively adapt the retrieval function of a meta-search engine to a particular group of users, outperforming Google in terms of retrieval quality after only a couple of hundred training examples.

Added

2026-05-08

Cumulated gain-based evaluation of IR techniques

Cumulated gain-based evaluation of IR techniques

Kalervo Järvelin, Jaana Kekäläinen

OrganizationsUniversity of Tampere

Why you should read this

Establishes the Normalized Discounted Cumulative Gain metric to accurately evaluate search results based on multiple levels of relevance and positional decay.

Modern large retrieval environments tend to overwhelm their users by their large output. Since all documents are not of equal relevance to their users, highly relevant documents should be identified and ranked first for presentation. In order to develop IR techniques in this direction, it is necessary to develop evaluation approaches and methods that credit IR methods for their ability to retrieve highly relevant documents. This can be done by extending traditional evaluation methods, that is, recall and precision based on binary relevance judgments, to graded relevance judgments. Alternatively, novel measures based on graded relevance judgments may be developed. This article proposes several novel measures that compute the cumulative gain the user obtains by examining the retrieval result up to a given ranked position. The first one accumulates the relevance scores of retrieved documents along the ranked result list. The second one is similar but applies a discount factor to the relevance scores in order to devaluate late-retrieved documents. The third one computes the relative-to-the-ideal performance of IR techniques, based on the cumulative gain they are able to yield. These novel measures are defined and discussed and their use is demonstrated in a case study using TREC data: sample system run results for 20 queries in TREC-7. As a relevance base we used novel graded relevance judgments on a four-point scale. The test results indicate that the proposed measures credit IR methods for their ability to retrieve highly relevant documents and allow testing of statistical significance of effectiveness differences. The graphs based on the measures also provide insight into the performance IR techniques and allow interpretation, for example, from the user point of view.

Added

2026-05-08