keyword
query expansion
Query expansion is an information retrieval technique that supplements an initial search query with additional relevant words, synonyms, or contextual phrases to improve search effectiveness. Its primary function is to resolve the vocabulary mismatch problem, which occurs when a searcher uses different terminology than the authors of relevant documents. Retrieval systems can generate expansion terms using global analysis of an entire corpus, local feedback derived from initially retrieved top documents, or neural language representations that identify semantically related vocabulary. By enriching the query beyond its original phrasing, query expansion helps search systems increase retrieval recall and rank relevant content that lacks the exact keywords of the initial request.
5 items

Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang
Why you should read this
Proposes a biomedical question-answering framework that boosts accuracy by using model-generated rationales for query formulation, balancing retrieval across diverse medical corpora, and filtering out distracting context with a perplexity-trained lightweight model.
Large language models (LLM) hold significant potential for applications in biomedicine, but they struggle with hallucinations and outdated knowledge. While retrieval-augmented generation (RAG) is generally employed to address these issues, it also has its own set of challenges: (1) LLMs are vulnerable to irrelevant or unhelpful context, (2) medical queries are often not well-targeted for helpful information, and (3) retrievers are prone to bias toward the specific source corpus they were trained on. In this study, we present RAG² (Rationale-Guided RAG), a new framework for enhancing the reliability of RAG in biomedical contexts. RAG² incorporates three key innovations: a small filtering model trained on perplexity-based labels of rationales, which selectively augments informative snippets of documents while filtering out distractors; LLM-generated rationales as queries to improve the utility of retrieved snippets; a structure designed to retrieve snippets evenly from a comprehensive set of four biomedical corpora, effectively mitigating retriever bias. Our experiments demonstrate that RAG² improves the state-of-the-art LLMs of varying sizes, with improvements of up to 6.1%, and it outperforms the previous best medical RAG model by up to 5.6% across three medical question-answering benchmarks. Our code is available at https://github.com/dmis-lab/RAG2
Added
2026-09-26

Query expansion using local and global document analysis
Jinxi Xu, W. Bruce Croft
Why you should read this
Introduces local context analysis, a query expansion technique combining global word-context features with locally retrieved passages to outperform traditional corpus-wide and local feedback retrieval methods.
Automatic query expansion has long been suggested as a technique for dealing with the fundamental issue of word mismatch in information retrieval. A number of approaches to expansion have been studied and, more recently, attention has focused on techniques that analyze the corpus to discover word relationships (global techniques) and those that analyze documents retrieved by the initial query ( local feedback). In this paper, we compare the effectiveness of these approaches and show that, although global analysis has some advantages, local analysis is generally more effective. We also show that using global analysis techniques, such as word context and phrase structure, on the local set of documents produces results that are both more effective and more predictable than simple local feedback.
Added
2026-09-25

IR evaluation methods for retrieving highly relevant documents
Kalervo Järvelin, Jaana Kekäläinen
Why you should read this
Introduces discounted cumulative gain (DCG) and cumulative gain metrics to evaluate information retrieval systems using graded, non-binary relevance judgments based on how effectively they prioritize highly relevant documents for users.
This paper proposes evaluation methods based on the use of non-dichotomous relevance judgements in IR experiments. It is argued that evaluation methods should credit IR methods for their ability to retrieve highly relevant documents. This is desirable from the user point of view in modern large IR enviroments. The proposed methods are (1) a novel application of P-R curves and average precision computations based on separate recall bases for documents of different degrees of relevance, and (2) two novel measures computing the cumulative gain the user obtains by examining the retrieval result up to a given ranked position. We then demonstrate the use of these evaluation methods in a case study on the effectiveness of query types, based on combinations of query structures and expansion, in retrieving documents of various degrees of relevance. The test was run with a best match retrieval system (InQuery¹) in a text database consisting of newspaper articles. The results indicate that the tested strong query structures are most effective in retrieving highly relevant documents. The differences between the query types are practically essential and statistically significant. More generally, the novel evaluation methods and the case demonstrate that non-dichotomous relevance assessments are applicable in IR experiments, may reveal interesting phenomena, and allow harder testing of IR methods.
Added
2026-09-25

SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking
Thibault Formal, Benjamin Piwowarski, Stéphane Clinchant
Why you should read this
Introduces SPLADE, a novel first-stage ranker that leverages explicit sparsity regularization and a log-saturation effect to achieve highly sparse representations and competitive results against state-of-the-art dense and sparse methods, while also exploring the effectiveness-efficiency trade-off.
In neural Information Retrieval, ongoing research is directed towards improving the first retriever in ranking pipelines. Learning dense embeddings to conduct retrieval using efficient approximate nearest neighbors methods has proven to work well. Meanwhile, there has been a growing interest in learning sparse representations for documents and queries, that could inherit from the desirable properties of bag-of-words models such as the exact matching of terms and the efficiency of inverted indexes. In this work, we present a new first-stage ranker based on explicit sparsity regularization and a log-saturation effect on term weights, leading to highly sparse representations and competitive results with respect to state-of-the-art dense and sparse methods. Our approach is simple, trained end-to-end in a single stage. We also explore the trade-off between effectiveness and efficiency, by controlling the contribution of the sparsity regularization.
Added
2026-05-26
License
Published with permission

The Probabilistic Relevance Framework: BM25 and Beyond
Stephen Robertson, Hugo Zaragoza
Why you should read this
Derives the BM25 scoring formula from fundamental probabilistic relevance principles to establish the mathematical justification for the most widely deployed baseline search algorithm.
The Probabilistic Relevance Framework (PRF) is a formal framework for document retrieval, grounded in work done in the 1970–1980s, which led to the development of one of the most successful text-retrieval algo¬rithms, BM25. In recent years, research in the PRF has yielded new retrieval models capable of taking into account document meta-data (especially structure and link-graph information). Again, this has led to one of the most successful Web-search and corporate-search algo¬rithms, BM25F. This work presents the PRF from a conceptual point of view, describing the probabilistic modelling assumptions behind the framework and the different ranking algorithms that result from its application: the binary independence model, relevance feedback mod¬els, BM25 and BM25F. It also discusses the relation between the PRF and other statistical models for IR, and covers some related topics, such as the use of non-textual features, and parameter optimisation for models with free parameters.
Added
2026-05-23
License
Published with permission
