LitSearch: A Retrieval Benchmark for Scientific Literature Search
Anirudh AjithMengzhou XiaAlexis ChevalierTanya GoyalDanqi ChenTianyu Gao
Introduces LitSearch, a high-quality retrieval benchmark of realistic scientific queries evaluated across full-text machine learning papers to expose critical shortcomings in current dense retrievers and commercial search engines.
Finding relevant scientific literature through specific natural language queries is essential for accelerating research and discovery. However, traditional citation recommendation benchmarks typically rely on raw inline citation contexts, which often produce noisy, trivial, or overly broad queries that do not mirror how researchers actually search. Modern information retrieval systems and commercial search engines frequently struggle to understand complex scientific concepts and retrieve relevant documents from extensive academic collections.
The article introduces LitSearch, a high-quality retrieval benchmark designed to evaluate how effectively modern search systems and language models retrieve scientific papers in response to realistic, natural research queries.
The researchers constructed a benchmark of 597 carefully curated literature search questions evaluated against a 64,183-paper corpus. The dataset comprises two subsets: 351 questions generated by prompting a large language model on citation mentions from published literature, and 246 questions authored directly by researchers about their recent conference publications. All questions underwent rigorous manual inspection and filtering for quality and specificity (categorized into broad and specific queries). The authors evaluated traditional keyword search, multiple modern dense retrieval models, and reranking pipelines using large language models, alongside a sample evaluation of commercial search tools.
The evaluation revealed that advanced dense retrieval models substantially outperform traditional keyword methods. Specifically, the best dense retriever (GritLM-7B) achieved a 74.8% recall within top-5 results on specific questions, outperforming traditional keyword search (BM25 at 50.0%) by an absolute 24.8 percentage points. Applying large language model reranking on top of dense retrieval further boosted top-5 recall to 79.2%. Conversely, commercial search engines struggled significantly, achieving a maximum top-5 recall of only 23.1% on inline-citation questions and trailing top retrieval models by up to 32 points. Furthermore, feeding full paper text into dense retrievers rather than just titles and abstracts did not consistently improve performance, often hindering retrieval due to context length constraints.
These findings indicate that general-purpose search engines and simple keyword matching are inadequate for professional academic literature discovery. Instruction-finetuned dense retrievers paired with language model rerankers provide a much more capable foundation for scientific discovery tools, offering significant potential to boost researcher productivity and reduce time spent on manual literature reviews. LitSearch also differentiates model performance more effectively than existing standard retrieval benchmarks.
Organizations developing scientific research tools should transition from legacy keyword search to instruction-tuned dense embedding architectures and integrate language model reranking pipelines. Tool builders should focus initial indexing on titles and abstracts until dense retrievers are better optimized for full-length scientific texts. Future work should focus on developing embedding models tailored to long-context document reasoning and expanding literature retrieval benchmarks across non-English languages and additional scientific domains.
The primary limitations include a scope restricted to English-language papers within natural language processing and machine learning, as well as author-written queries that exhibit higher term overlap and are easier than inline questions. Nevertheless, the manual validation across hundreds of peer-reviewed papers supports high confidence in LitSearch as an informative and challenging standard for scientific information retrieval.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). BEIR establishes the broad zero-shot retrieval benchmark tradition and model comparisons that provide essential context for LitSearch’s evaluation design.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). ANCE explains hard-negative training as a key foundation for dense retrievers whose performance LitSearch compares against keyword search.
- Paper: Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval, Luyu Gao et al. (2022). coCondenser shows how corpus-aware pretraining can strengthen dense retrieval, clarifying the model family behind LitSearch’s retrieval results.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). Passage Re-ranking with BERT introduces the neural reranking approach that helps explain LitSearch’s gains from reranking retrieved candidates.
- Paper: IR evaluation methods for retrieving highly relevant documents, Kalervo Järvelin et al. (2000). This paper introduces graded relevance and discounted cumulative gain, concepts needed to understand how ranked retrieval results are evaluated.
- Paper: SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval, Hossein A. Rahmani et al. (2025). SynDL extends retrieval benchmarking beyond LitSearch by testing whether large-scale synthetic relevance judgments can make system evaluation broader and less costly.
