BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers
Ran XuWenqi ShiYue YuYuchen ZhuangYanqiao ZhuMay Dongmei WangJoyce C. HoChao ZhangCarl Yang
Presents BMRetriever, an open-source family of dense text retrievers that combines unsupervised contrastive pre-training with synthetic instruction tuning to outperform much larger biomedical models across 11 standard benchmarks within academic compute budgets.
Reliable information retrieval in biomedicine is critical for supporting applications such as clinical decision-making, medical question answering, and scientific discovery. However, developing high-performing biomedical text retrievers has historically been hindered by the scarcity of publicly annotated domain data and the high computational expense of training large language models. Existing specialized models often depend on small architectures or inaccessible proprietary datasets, while general-domain models frequently struggle when applied to specialized medical terminology.
The article introduces and evaluates BMRetriever, a series of open-access dense text retrieval models spanning 410 million to 7 billion parameters. The primary objective is to demonstrate that large language models can be efficiently adapted into specialized biomedical retrievers using only publicly available data and manageable computational resources.
The developers designed a two-stage training approach. First, the models undergo unsupervised contrastive pre-training on large public biomedical corpora, including scientific papers and medical textbooks, to learn domain-specific terminology. Second, the models undergo multi-task instruction fine-tuning using a mix of human-annotated datasets and synthetic query-passage pairs generated by artificial intelligence models. The framework was evaluated across 11 benchmark datasets spanning five core biomedical tasks: standard information retrieval, sentence similarity, question answering, entity linking, and paper recommendation.
The evaluation produced several key findings. First, BMRetriever achieves state-of-the-art retrieval accuracy while demonstrating exceptional parameter efficiency; the compact 410-million-parameter version outperformed competing baseline models that were up to 11.7 times larger. Second, the intermediate 1-billion and 2-billion parameter variants matched or exceeded the performance of models with more than 5 billion parameters, achieving over 98% of the top-performing 7-billion model's accuracy. Third, synthetic fine-tuning data contributed the largest single gain in model adaptability and task generalization among the training sources. Finally, training remained cost-effective, requiring only 10 million pre-training pairs and generating synthetic data for under $500 in total computing costs.
These results demonstrate that organizations can deploy highly accurate, domain-specific retrieval systems without relying on massive compute budgets or sensitive proprietary user data. By achieving high accuracy at smaller model sizes, the approach significantly reduces hardware, hosting, and operational costs. The open availability of the training recipe and model checkpoints also enhances compliance transparency and simplifies adoption for technical teams.
Decision-makers should consider adopting the 1-billion or 2-billion parameter BMRetriever models for biomedical search pipelines, as they provide the best balance of retrieval accuracy and computational efficiency. Future efforts should focus on optimizing inference speed and embedding storage costs, as larger model architectures introduce greater latency overhead compared to legacy, smaller encoder systems.
Confidence in the reported benchmarks is high due to rigorous evaluations across diverse standard datasets and manual medical review verifying that synthetic training data did not introduce factual errors. However, users should remain mindful of latency trade-offs during real-time deployment and validate model behavior against specialized local clinical vocabularies.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). Contriever introduces the foundational contrastive pre-training methods for unsupervised dense information retrieval that BMRetriever adapts to the biomedical domain.
- Paper: Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing, Yu Gu et al. (2020). This paper establishes PubMedBERT and domain-specific pre-training on PubMed corpora, demonstrating the necessity of specialized pre-training that BMRetriever adopts.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). BioBERT pioneered pre-training contextual representations on large biomedical corpora for downstream medical text mining and question answering.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). BEIR defines the standard zero-shot retrieval evaluation benchmark and task paradigm used to assess dense retrieval performance.
- Paper: Precise Zero-Shot Dense Retrieval without Relevance Labels, Luyu Gao et al. (2023). HyDE demonstrates how LLM-generated synthetic text can enhance zero-shot dense retrieval, providing context for BMRetriever's use of synthetic pairs.
- Paper: PubMedQA: A Dataset for Biomedical Research Question Answering, Qiao Jin et al. (2019). PubMedQA provides a key biomedical question answering benchmark widely used to evaluate domain-specific retrieval-augmented pipelines.
- Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams, Di Jin et al. (2020). MedQA introduces the standard multi-choice clinical exam benchmark for assessing evidence retrieval and biomedical reasoning.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). ANCE formalizes hard negative mining techniques in contrastive learning that are essential for effectively training dense text retrievers.
- Paper: Rationale-Guided Retrieval Augmented Generation for Medical Question Answering, Jiwoong Sohn et al. (2025). RAG2 builds on domain-specific biomedical retrieval by using rationale-guided querying and multi-corpus evidence filtering for clinical question answering.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG extends dense biomedical retrieval pipelines by instruction-tuning LLMs to unify context reranking and answer generation.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). NV-Embed explores advanced instruction-tuning and attention architectures to scale generalist decoder-only embedding models beyond specialized retrieval domains.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). Qwen3 Embedding advances foundation-model text embeddings and reranking using multi-stage synthetic data pipelines across extensive sequence lengths.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). bge-en-icl generalizes instruction-tuned text embedders by incorporating in-context learning to adapt to novel retrieval tasks via few-shot demonstration queries.
