Query expansion using local and global document analysis
Jinxi XuW. Bruce Croft
Introduces local context analysis, a query expansion technique combining global word-context features with locally retrieved passages to outperform traditional corpus-wide and local feedback retrieval methods.
Information retrieval systems frequently fail because users submit short queries using different vocabulary than the authors of relevant documents. Automatic query expansion solves this word mismatch problem by enriching queries with related terms. However, traditional techniques face severe trade-offs: analyzing an entire text repository is computationally expensive and introduces off-topic terms, while standard local feedback based on initial search results performs unpredictably when early matches contain non-relevant documents.
The article evaluates existing query expansion techniques and introduces a hybrid method called local context analysis, which applies structural and context-aware term selection to top-ranked passages rather than entire documents or full collections. The authors conducted comparative experiments across three major benchmark collections—TREC-3, TREC-4, and the legal WEST database—evaluating overall precision, recall, and query-by-query reliability across hundreds of thousands of documents.
The analysis demonstrates that local context analysis consistently outperforms both global collection-level analysis and traditional local feedback. On large collections, local context analysis improved average retrieval precision by roughly 23% to 24% over unexpanded baseline searches, compared to modest gains of 3% to 8% for global analysis and 14% to 21% for local feedback. Furthermore, local context analysis proved substantially more robust on difficult queries: on the TREC-4 set, it improved 38 of 49 queries and harmed only 11, whereas standard local feedback degraded 21 queries—frequently destroying performance on queries that started with low initial precision. Finally, local context analysis showed low sensitivity to parameter tuning, maintaining high effectiveness across a broad window of 30 to 300 retrieved passages.
These findings indicate that search systems can achieve significant gains in retrieval accuracy and reliability without the massive computational and storage overhead required to build corpus-wide concept databases. By extracting noun phrases from focused 300-word passages and enforcing term co-occurrence constraints across all original query words, local context analysis provides an efficient mechanism suitable for interactive, real-time search environments.
Organizations implementing automated search expansion should adopt passage-level local context analysis over whole-document local feedback, especially when serving domains where individual search failure carries high operational risk. When applying expansion to specialized databases with high baseline accuracy and fewer relevant documents, systems should downweight expansion terms to prevent query drift. Future technical work should focus on developing adaptive algorithms that automatically calibrate passage counts and term weights for each query rather than relying on static system parameters.
- Paper: Term-weighting approaches in automatic text retrieval, Gerard Salton et al. (1988). Establishes the fundamental single-term tf-idf weighting and document length normalization frameworks that form the foundation of vector-space information retrieval and term expansion.
- Paper: Indexing By Latent Semantic Analysis, Scott Deerwester et al. (1990). Introduces Latent Semantic Indexing to address vocabulary mismatch by extracting global term-document co-occurrence structures, serving as a primary baseline for global analysis techniques.
- Paper: Some simple effective approximations to the 2-Poisson model for probabilistic weighted retrieval, Stephen Robertson et al. (1994). Provides the 2-Poisson probabilistic weighting foundations (such as the BM series) that underpin modern scoring and relevance feedback approximations in information retrieval.
- Paper: Scatter/Gather: a cluster-based approach to browsing large document collections, Douglass R. Cutting et al. (1992). Pioneers dynamic document grouping and cluster-based browsing over retrieved subsets, laying groundwork for analyzing local document contexts.
- Paper: Relevance-Based Language Models, Victor Lavrenko et al. (2001). Extends statistical feedback and local document analysis into the language modeling framework through formal relevance-based language models.
- Paper: A study of smoothing methods for language models applied to Ad Hoc information retrieval, ChengXiang Zhai et al. (2001). Analyzes language modeling smoothing methods that mathematically refine term-weighting and query generation beyond heuristic local and global feedback techniques.
- Paper: The Probabilistic Relevance Framework: BM25 and Beyond, Stephen Robertson et al. (2009). Synthesizes probabilistic relevance feedback, pseudo-relevance feedback, and BM25 term weighting into a comprehensive theoretical framework.
- Paper: Automatic Retrieval and Clustering of Similar Words, Dekang Lin (1998). Develops formal information-theoretic distributional similarity measures from text corpora to automate thesaurus construction and vocabulary expansion.
- Paper: Mining the Web for Synonyms: PMI-IR versus LSA on TOEFL, Peter D. Turney (2001). Demonstrates how large-scale corpus statistics and search queries can be used for global synonym discovery via pointwise mutual information.
- Paper: Probabilistic Latent Semantic Analysis, Thomas Hofmann (1999). Provides a sound generative probabilistic formulation (PLSA) for modeling global latent word associations that improves upon algebraic dimensionality reduction.
- Paper: Automatic image annotation and retrieval using cross-media relevance models, J. Jeon et al. (2003). Generalizes relevance feedback and local co-occurrence models to cross-modal tasks by jointly estimating text queries and visual features.
- Paper: CeQe: Grounding Lexical Retrieval in Semantic Evidence, Adam Kahirov et al.. Applies modern cross-encoder and neural semantic representations to conduct query expansion for hybrid BM25 retrieval pipelines.
- Paper: SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, Thibault Formal et al. (2021). Modernizes sparse term expansion by learning contextual term weights and vocabulary expansions end-to-end using deep language models.
