LitSearch: A Retrieval Benchmark for Scientific Literature Search

Anirudh AjithMengzhou XiaAlexis ChevalierTanya GoyalDanqi ChenTianyu Gao

article2024EMNLP68 citations

Introduces LitSearch, a high-quality retrieval benchmark of realistic scientific queries evaluated across full-text machine learning papers to expose critical shortcomings in current dense retrievers and commercial search engines.

Listen

Finding relevant scientific literature through specific natural language queries is essential for accelerating research and discovery. However, traditional citation recommendation benchmarks typically rely on raw inline citation contexts, which often produce noisy, trivial, or overly broad queries that do not mirror how researchers actually search. Modern information retrieval systems and commercial search engines frequently struggle to understand complex scientific concepts and retrieve relevant documents from extensive academic collections.

The article introduces LitSearch, a high-quality retrieval benchmark designed to evaluate how effectively modern search systems and language models retrieve scientific papers in response to realistic, natural research queries.

The researchers constructed a benchmark of 597 carefully curated literature search questions evaluated against a 64,183-paper corpus. The dataset comprises two subsets: 351 questions generated by prompting a large language model on citation mentions from published literature, and 246 questions authored directly by researchers about their recent conference publications. All questions underwent rigorous manual inspection and filtering for quality and specificity (categorized into broad and specific queries). The authors evaluated traditional keyword search, multiple modern dense retrieval models, and reranking pipelines using large language models, alongside a sample evaluation of commercial search tools.

The evaluation revealed that advanced dense retrieval models substantially outperform traditional keyword methods. Specifically, the best dense retriever (GritLM-7B) achieved a 74.8% recall within top-5 results on specific questions, outperforming traditional keyword search (BM25 at 50.0%) by an absolute 24.8 percentage points. Applying large language model reranking on top of dense retrieval further boosted top-5 recall to 79.2%. Conversely, commercial search engines struggled significantly, achieving a maximum top-5 recall of only 23.1% on inline-citation questions and trailing top retrieval models by up to 32 points. Furthermore, feeding full paper text into dense retrievers rather than just titles and abstracts did not consistently improve performance, often hindering retrieval due to context length constraints.

These findings indicate that general-purpose search engines and simple keyword matching are inadequate for professional academic literature discovery. Instruction-finetuned dense retrievers paired with language model rerankers provide a much more capable foundation for scientific discovery tools, offering significant potential to boost researcher productivity and reduce time spent on manual literature reviews. LitSearch also differentiates model performance more effectively than existing standard retrieval benchmarks.

Organizations developing scientific research tools should transition from legacy keyword search to instruction-tuned dense embedding architectures and integrate language model reranking pipelines. Tool builders should focus initial indexing on titles and abstracts until dense retrievers are better optimized for full-length scientific texts. Future work should focus on developing embedding models tailored to long-context document reasoning and expanding literature retrieval benchmarks across non-English languages and additional scientific domains.

The primary limitations include a scope restricted to English-language papers within natural language processing and machine learning, as well as author-written queries that exhibit higher term overlap and are easier than inline questions. Nevertheless, the manual validation across hundreds of peer-reviewed papers supports high confidence in LitSearch as an informative and challenging standard for scientific information retrieval.

Cover for LitSearch: A Retrieval Benchmark for Scientific Literature Search

Abstract

Literature search questions, such as “Where can I find research on the evaluation of consistency in generated summaries?” pose significant challenges for modern search engines and retrieval systems. These questions often require a deep understanding of research concepts and the ability to reason across entire articles. In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers. All LitSearch questions were manually examined or edited by experts to ensure high quality. We extensively benchmark state-of-the-art retrieval models and also evaluate two LLM-based reranking pipelines. We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in recall@5. The LLM-based reranking strategies further improve the best-performing dense retriever by 4.4%. Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points. Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case.

Table of Contents

  • 1 Introduction
  • 2 LitSearch
  • 2.1 Inline-citation Questions
  • 2.2 Author-written Questions
  • 2.3 Manual Filtering to Ensure High Quality
  • 2.4 Dataset Statistics
  • 2.5 The Retrieval Corpus
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Baselines
  • 3.3 Results
  • 4 Analysis
  • 4.1 Does Including More Paper Content Improve Retrieval Performance?
  • 4.2 Does the Source of Inline Citation Questions Matter?
  • 4.3 Performance of Search Engines
  • 4.4 Comparing Other Retrieval Benchmarks
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Annotator Acknowledgments
  • B Annotation Details
  • C Retrieval Corpus
  • D Retriever details
  • E Prompts and Additional Statistics

Knowls

  1. Knowl 1 — LitSearch benchmarks realistic scientific literature search

    definition

    LitSearch is a retrieval benchmark for natural-language questions that researchers might ask when looking for scientific literature. Each of its 597 questions is paired with one or more relevant target papers. The benchmark contains 351 questions derived from inline citation contexts and 246 questions written by authors about their own papers. It is designed to test retrieval of relevant papers from a scientific corpus, rather than retrieval from a short passage collection.

  2. Knowl 2 — Question collection combines citation-context generation with expert author questions

    model/method

    LitSearch collects questions through two complementary routes. For inline-citation questions, the authors sampled citation mentions from the Semantic Scholar Open Research Corpus (S2ORC, 2024-03-26 version); target papers were restricted to the ACL Anthology, while source papers could come from any venue. GPT-4 received the citation paragraph and cited-paper titles, along with two demonstrations, and was asked to formulate a literature-search question answerable by the cited papers. The authors filtered questions whose word overlap with the target-paper title exceeded 0.3 for ACL-sourced questions or 0.1 for non-ACL-sourced questions; this removed 5% and 80% of those questions, respectively.

    For author-written questions, researchers were invited to write a literature-search question about their own ACL 2023 or ICLR 2024 paper. Invitations went to 623 ACL authors and 404 ICLR authors; 175 and 117 questions were received, respectively. The retained counts were 155 ACL questions and 91 ICLR questions. The benchmark authors manually examined every question, assigned quality and specificity labels, excluded quality-0 questions, and made minor corrections where needed. For inline-citation questions, 98 of 382 reviewed ACL-sourced questions and 253 of 450 reviewed non-ACL-sourced questions were retained.

  3. Knowl 3 — Manual labels define question specificity and quality

    definition

    LitSearch labels questions along two dimensions. A question is broad when up to 20 papers in the corpus could fit it, and specific when up to 5 papers could fit it. The quality scale assigns 0 to questions that are factually wrong, unrealistic, too broad or too specific, or too easy; 1 to acceptable questions that may be somewhat out of distribution or relatively easy because of title or abstract overlap; and 2 to meaningful, challenging questions. Only quality-1 and quality-2 questions are included. These labels allow retrieval performance to be reported separately for questions with different expected difficulty and breadth.

  4. Knowl 4 — Corpus and query subsets provide a large, varied retrieval testbed

    data/table

    The retrieval corpus contains ACL Anthology and ICLR papers extracted from S2ORC: 59,383 ACL Anthology entries and 4,807 ICLR entries, with 7 papers shared across the subsets, for 64,183 unique papers. Average document length is 134 words for titles and abstracts and 6,041 words for full text. The following statistics describe the question subsets; overlap is the fraction of question words also appearing in target-paper titles and abstracts, and average target papers counts the annotated relevant papers, not all corpus papers that might satisfy a query.

    Question source Specificity Questions Avg. length (words) Overlap Avg. target papers
    Inline-citation Broad 120 20.6 0.33 1.21
    Inline-citation Specific 231 22.1 0.34 1.07
    Author-written Broad 35 15.8 0.43 1.03
    Author-written Specific 211 17.9 0.43 1.00

    Author-written questions have greater title-and-abstract overlap than inline-citation questions (0.43 versus 0.33–0.34), consistent with authors reusing terminology from their own papers.

  5. Knowl 5 — Retrieval evaluation compares lexical and dense retrievers on paper abstracts

    experimental setup

    The default retrieval documents are paper titles and abstracts, averaging 134 words, because of embedding-model context limits. The evaluated systems are BM25 and four dense retrievers: GTR-T5-large, Instructor-XL, E5-large-v2, and GritLM-7B. For broad questions, performance is measured by recall@20; for specific questions, it is measured by recall@5 and recall@20. The paper reports results separately for inline-citation and author-written questions as well as averages across these subsets. When testing full-text retrieval, the model context limits are 512 tokens for GTR-T5-large, Instructor-XL, and E5-large-v2, and 2,048 tokens for GritLM-7B.

  6. Knowl 6 — GPT-4o reranking uses either retrieved papers alone or citation-neighbor expansion

    model/method

    The vanilla reranking pipeline sends the titles and abstracts of the top 100 papers from a retriever to GPT-4o, which ranks them by relevance to the search question. Its average prompt context is 13,844 words. The one-hop pipeline starts with the top 50 retrieved papers, adds papers cited by each seed paper, removes duplicates, and orders candidates by alternating each seed paper with its cited papers. It truncates this candidate list to 200 papers and applies the same GPT-4o ranking procedure; its average context is 27,544 words. The one-hop design tests whether citation links can help recover relevant targets from initially retrieved papers.

  7. Knowl 7 — GritLM-7B is the strongest base retriever and GPT-4o reranking improves it

    empirical result

    Using titles and abstracts for retrieval and reranking, GritLM-7B achieves the strongest base-retrieval averages: 70.8% recall@20 on broad questions and 74.8% recall@5 on specific questions. GPT-4o vanilla reranking with GritLM raises these averages to 75.3% and 79.2%, respectively; one-hop reranking yields 73.2% and 77.0%. The full results below report recall percentages by question source and specificity. For specific questions, both recall@5 and recall@20 are shown; broad-question results are recall@20.

    System Inline broad R@20 Inline specific R@5 Inline specific R@20 Author broad R@20 Author specific R@5 Author specific R@20 Avg. broad R@20 Avg. specific R@5
    BM25 37.4 38.5 55.8 48.6 62.6 73.5 39.9 50.0
    GTR-T5-large 45.7 38.5 51.5 37.1 40.8 55.9 43.8 39.6
    Instructor-XL 56.3 48.9 60.0 57.1 55.9 70.1 56.5 52.3
    E5-large-v2 55.8 50.4 63.9 54.3 62.6 75.8 55.4 56.2
    GritLM-7B 69.7 67.7 77.9 74.3 82.5 89.1 70.8 74.8
    GPT-4o reranking with BM25 54.9 60.0 67.5 77.1 76.8 82.9 59.9 68.0
    GPT-4o one-hop with BM25 62.0 64.1 71.6 74.3 73.5 77.7 64.8 68.6
    GPT-4o reranking with GritLM 74.7 73.2 79.9 77.1 85.8 92.4 75.3 79.2
    GPT-4o one-hop with GritLM 72.9 70.3 78.4 74.3 84.4 87.2 73.2 77.0

    All instruction-finetuned dense retrievers outperform BM25 on the aggregate measures, and GritLM-7B is the best dense retriever. Its average recall@5 advantage over BM25 on specific questions is 24.8 percentage points; GPT-4o vanilla reranking adds 4.4 points to GritLM's average recall@5.

  8. Knowl 8 — Retrieval difficulty varies with question quality, source, and specificity

    empirical result

    Questions judged quality 2—more realistic and challenging—have lower recall@5 than quality-1 questions for every tested retriever on both inline-citation and author-written specific subsets:

    Retriever Inline quality 1 R@5 Inline quality 2 R@5 Author quality 1 R@5 Author quality 2 R@5
    BM25 36.4 30.6 62.2 55.0
    GTR-T5-large 42.0 31.4 40.7 36.9
    Instructor-XL 55.1 39.5 58.5 48.6
    E5-large-v2 48.4 42.6 61.5 57.7
    GritLM-7B 67.3 58.7 80.0 76.6

    Across systems, author-written questions are generally easier than inline-citation questions; the authors attribute this partly to their higher overlap with paper titles and abstracts. Recall is generally higher on specific than broad questions, which have more competing papers. For BM25 on specific inline-citation questions, one-hop reranking improves on vanilla reranking by 11.4 recall points for ACL-sourced questions (71.2% versus 59.8%), but only 1.2 points for non-ACL-sourced questions (61.2% versus 60.0%). This pattern is consistent with BM25 exploiting the citation-source relationship when the source paper is in the retrieval corpus. GritLM does not show the same large one-hop advantage.

  9. Knowl 9 — Adding full paper text does not consistently improve retrieval

    empirical result

    The authors compared retrieval over titles and abstracts with retrieval over full paper text, truncated to each model's context limit. Full papers average 6,041 words, far longer than the models' 512- or 2,048-token limits. The table reports recall percentages for broad questions at R@20 and specific questions at R@5. Including full text produces no consistent gain: it improves some author-written broad results, especially for BM25, but lowers performance for many other model-and-subset combinations, including GritLM on both author-written metrics.

    Retriever Document input Inline broad R@20 Inline specific R@5 Author broad R@20 Author specific R@5
    BM25 Title and abstract 37.4 38.5 48.6 62.6
    BM25 Full text 18.6 23.8 65.7 71.6
    GTR-T5-large Title and abstract 45.7 38.5 37.1 40.8
    GTR-T5-large Full text 43.9 39.4 45.7 39.8
    Instructor-XL Title and abstract 56.3 48.9 57.1 55.9
    Instructor-XL Full text 53.0 50.9 57.1 56.9
    E5-large-v2 Title and abstract 55.8 50.4 54.3 62.6
    E5-large-v2 Full text 56.9 48.7 60.0 62.1
    GritLM-7B Title and abstract 69.7 67.7 74.3 82.5
    GritLM-7B Full text 70.8 63.4 65.7 73.0
  10. Knowl 10 — Commercial search tools perform unevenly on specific questions

    empirical result

    A human study tested Google Search, Google Scholar, and Elicit on a random sample of 80 specific questions: 40 inline-citation and 40 author-written. The reported metric is recall@5. For Google Search, only academic papers among the first page of results were considered, up to five papers. Search-engine results are not directly comparable with the benchmark retrievers because the commercial tools search a larger corpus.

    System Inline-citation R@5 Author-written R@5
    BM25 38.5 62.6
    GritLM-7B 67.7 82.5
    Google Search 23.1 62.5
    Google Scholar 20.5 17.5
    Elicit 23.1 17.5

    All three commercial tools have low recall on inline-citation questions. Google Search performs substantially better than Google Scholar and Elicit on author-written questions, but remains below GritLM-7B.

  11. Knowl 11 — LitSearch differentiates embedding models that score similarly on other benchmarks

    empirical result

    The authors compared embedding-model performance across MS MARCO, SCIDOCS, Natural Questions (NQ), ArXiv, and LitSearch. Every value below is nDCG@10, enabling a common metric comparison. LitSearch broadly agrees with the other benchmarks but separates model performance more clearly in some cases: on specific LitSearch questions, GritLM-7B scores 60.3 versus 45.3 for E5-large-v2, a 15-point gap, whereas their MS MARCO scores are close at 42.0 and 43.5.

    Retriever MS MARCO SCIDOCS NQ ArXiv LitSearch broad LitSearch specific
    GTR-T5-large 42.7 15.5 55.1 17.5 23.3 30.4
    Instructor-XL 41.6 17.4 57.2 19.8 32.8 41.2
    E5-large-v2 43.5 20.5 63.4 27.0 27.1 45.3
    GritLM-7B 42.0 24.4 70.3 34.3 44.1 60.3
  12. Knowl 12 — Benchmark limitations include coverage, curation, and evaluation constraints

    limitation

    The authors note that manual review does not eliminate questions that are somewhat unlike researchers' usual queries or are easy because of overlap with target papers. Author-written questions were easier than expected, since writing challenging literature-search questions proved difficult even for experienced researchers. The evaluated retrievers and rerankers are not an exhaustive set of possible systems, and the benchmark focuses on English questions and papers. The authors also identify potential dataset biases: inline-citation sampling may overrepresent highly cited target papers, while author-written questions cover only ACL 2023 and ICLR 2024 papers.

Coverage note — Supplementary verbatim prompt templates and the long annotator roster are omitted because they document implementation and contributors rather than add distinct scientific findings.

References

  1. 1.Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. Llm2vec: Large language models are secretly powerful text encoders. Preprint, arXiv:2404.05961.
  2. 2.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China. Association for Computational Linguistics.
  3. 3.Chandra Bhagavatula, Sergey Feldman, Russell Power, and Waleed Ammar. 2018. Content-based citation recommendation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 238–251, New Orleans, Louisiana. Association for Computational Linguistics.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  5. 5.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020. SPECTER: Document-level representation learning using citation-informed transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, Online. Association for Computational Linguistics.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional Transformers for language understanding. In North American Chapter of the Association for Computational Linguistics (NAACL).
  7. 7.Michael Färber and Adam Jatowt. 2020. Citation recommendation: approaches and datasets. Int. J. Digit. Libr., 21(4):375–405.
  8. 8.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Empirical Methods in Natural Language Processing (EMNLP), pages 6894–6910.
  9. 9.Nianlong Gu, Yingqiang Gao, and Richard H. R. Hahnloser. 2022. Local citation recommendation with hierarchical-attention text encoder and scibert-based reranking. In Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10–14, 2022, Proceedings, Part I, page 274–288, Berlin, Heidelberg. Springer-Verlag.
  10. 10.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  11. 11.Qi He, Jian Pei, Daniel Kifer, Prasenjit Mitra, and Lee Giles. 2010. Context-aware citation recommendation. In Proceedings of the 19th International Conference on World Wide Web, WWW ’10, page 421–430, New York, NY, USA. Association for Computing Machinery.
  12. 12.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  13. 13.Chanwoo Jeong, Sion Jang, Eunjeong Park, and Sungchul Choi. 2020. A context-aware citation recommendation model with bert and graph convolutional networks. Scientometrics, 124(3):1907–1922.
  14. 14.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  15. 15.O. Khattab and Matei A. Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval.
  16. 16.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  17. 17.Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. Nv-embed: Improved techniques for training llms as generalist embedding models. Preprint, arXiv:2405.17428.
  18. 18.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Association for Computational Linguistics (ACL), pages 6086–6096.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  20. 20.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Association for Computational Linguistics (ACL), pages 4969–4983.
  21. 21.Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023. Zero-shot listwise document reranking with a large language model. arXiv preprint arXiv:2305.02156.
  22. 22.Zoran Medic and Jan Snajder. 2020. Improved local citation recommendation based on context enhanced with global information. In Proceedings of the First Workshop on Scholarly Document Processing, pages 97–103, Online. Association for Computational Linguistics.
  23. 23.Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. 2024. Generative representational instruction tuning. arXiv preprint arXiv:2402.09906.
  24. 24.Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022. Mteb: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316.
  25. 25.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2017. MS MARCO: A human-generated MAchine reading COmprehension dataset.
  26. 26.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large dual encoders are generalizable retrievers. In Empirical Methods in Natural Language Processing (EMNLP), pages 9844–9855.
  27. 27.OpenAI. 2023. GPT-4 Technical Report. Preprint, arXiv:2303.08774.
  28. 28.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543.
  29. 29.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  30. 30.Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. 2024. Citeme: Can language models accurately cite scientific claims? Preprint, arXiv:2407.12861.
  31. 31.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Empirical Methods in Natural Language Processing and International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).
  32. 32.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  33. 33.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. One embedder, any task: Instruction-finetuned text embeddings. In Findings of the Association for Computational Linguistics: ACL 2023, pages 1102–1121, Toronto, Canada. Association for Computational Linguistics.
  34. 34.Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023a. Is chatgpt good at search? investigating large language models as re-ranking agent. ArXiv, abs/2304.09542.
  35. 35.Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023b. Is ChatGPT good at search? investigating large language models as re-ranking agents. In Empirical Methods in Natural Language Processing (EMNLP), pages 14918–14937.
  36. 36.Michael Tang, Shunyu Yao, John Yang, and Karthik Narasimhan. 2023. Referral augmentation for zero-shot information retrieval. Preprint, arXiv:2305.15098.
  37. 37.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  38. 38.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971.
  39. 39.Ellen M. Voorhees and Dawn M. Tice. 2000. Building a question answering test collection. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’00, page 200–207, New York, NY, USA. Association for Computing Machinery.
  40. 40.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training. ArXiv, abs/2212.03533.
  41. 41.Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang. 2022. Grit: A generative region-to-text transformer for object understanding. Preprint, arXiv:2212.00280.

Citation

MLA
Ajith, A., et al. “LitSearch: A Retrieval Benchmark for Scientific Literature Search”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 15068–83, https://doi.org/10.18653/v1/2024.emnlp-main.840.
APA
Ajith, A., Xia, M., Chevalier, A., Goyal, T., Chen, D., & Gao, T. (2024). LitSearch: A Retrieval Benchmark for Scientific Literature Search. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 15068–15083. https://doi.org/10.18653/v1/2024.emnlp-main.840
Chicago
Ajith, A., M. Xia, A. Chevalier, T. Goyal, D. Chen, and T. Gao. 2024. “LitSearch: A Retrieval Benchmark for Scientific Literature Search”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 15068–83. https://doi.org/10.18653/v1/2024.emnlp-main.840.
Harvard
Ajith, A. et al. (2024) “LitSearch: A Retrieval Benchmark for Scientific Literature Search”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 15068–15083. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.840.
Vancouver
1. Ajith A, Xia M, Chevalier A, Goyal T, Chen D, Gao T (2024) LitSearch: A Retrieval Benchmark for Scientific Literature Search. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 15068–15083

BibTeX

@inproceedings{ajith-etal-2024-litsearch,
    title = "{L}it{S}earch: A Retrieval Benchmark for Scientific Literature Search",
    author = "Ajith, Anirudh  and
      Xia, Mengzhou  and
      Chevalier, Alexis  and
      Goyal, Tanya  and
      Chen, Danqi  and
      Gao, Tianyu",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.840/",
    doi = "10.18653/v1/2024.emnlp-main.840",
    pages = "15068--15083"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/