BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers

Ran XuWenqi ShiYue YuYuchen ZhuangYanqiao ZhuMay Dongmei WangJoyce C. HoChao ZhangCarl Yang

article2024EMNLP66 citations

Presents BMRetriever, an open-source family of dense text retrievers that combines unsupervised contrastive pre-training with synthetic instruction tuning to outperform much larger biomedical models across 11 standard benchmarks within academic compute budgets.

Listen

Reliable information retrieval in biomedicine is critical for supporting applications such as clinical decision-making, medical question answering, and scientific discovery. However, developing high-performing biomedical text retrievers has historically been hindered by the scarcity of publicly annotated domain data and the high computational expense of training large language models. Existing specialized models often depend on small architectures or inaccessible proprietary datasets, while general-domain models frequently struggle when applied to specialized medical terminology.

The article introduces and evaluates BMRetriever, a series of open-access dense text retrieval models spanning 410 million to 7 billion parameters. The primary objective is to demonstrate that large language models can be efficiently adapted into specialized biomedical retrievers using only publicly available data and manageable computational resources.

The developers designed a two-stage training approach. First, the models undergo unsupervised contrastive pre-training on large public biomedical corpora, including scientific papers and medical textbooks, to learn domain-specific terminology. Second, the models undergo multi-task instruction fine-tuning using a mix of human-annotated datasets and synthetic query-passage pairs generated by artificial intelligence models. The framework was evaluated across 11 benchmark datasets spanning five core biomedical tasks: standard information retrieval, sentence similarity, question answering, entity linking, and paper recommendation.

The evaluation produced several key findings. First, BMRetriever achieves state-of-the-art retrieval accuracy while demonstrating exceptional parameter efficiency; the compact 410-million-parameter version outperformed competing baseline models that were up to 11.7 times larger. Second, the intermediate 1-billion and 2-billion parameter variants matched or exceeded the performance of models with more than 5 billion parameters, achieving over 98% of the top-performing 7-billion model's accuracy. Third, synthetic fine-tuning data contributed the largest single gain in model adaptability and task generalization among the training sources. Finally, training remained cost-effective, requiring only 10 million pre-training pairs and generating synthetic data for under $500 in total computing costs.

These results demonstrate that organizations can deploy highly accurate, domain-specific retrieval systems without relying on massive compute budgets or sensitive proprietary user data. By achieving high accuracy at smaller model sizes, the approach significantly reduces hardware, hosting, and operational costs. The open availability of the training recipe and model checkpoints also enhances compliance transparency and simplifies adoption for technical teams.

Decision-makers should consider adopting the 1-billion or 2-billion parameter BMRetriever models for biomedical search pipelines, as they provide the best balance of retrieval accuracy and computational efficiency. Future efforts should focus on optimizing inference speed and embedding storage costs, as larger model architectures introduce greater latency overhead compared to legacy, smaller encoder systems.

Confidence in the reported benchmarks is high due to rigorous evaluations across diverse standard datasets and manual medical review verifying that synthetic training data did not introduce factual errors. However, users should remain mindful of latency trade-offs during real-time deployment and validate model behavior against specialized local clinical vocabularies.

Xu et al (2024).pdf
Cover for BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers

Abstract

Developing effective biomedical retrieval models is important for excelling at knowledge-intensive biomedical tasks but still challenging due to the lack of sufficient publicly annotated biomedical data and computational resources. We present BMRetriever, a series of dense retrievers for enhancing biomedical retrieval via unsupervised pre-training on large biomedical corpora, followed by instruction fine-tuning on a combination of labeled datasets and synthetic pairs. Experiments on 5 biomedical tasks across 11 datasets verify BMRetriever's efficacy on various biomedical applications. BMRetriever also exhibits strong parameter efficiency, with the 410M variant outperforming baselines up to 11.7 times larger, and the 2B variant matching the performance of models with over 5B parameters. The training data and model checkpoints are released at https://huggingface.co/BMRetriever to ensure transparency, reproducibility, and application to new domains.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Background of Dense Text Retrieval
  • 3.2 Unsupervised Contrastive Pre-training
  • 3.3 Supervised Instruction Fine-tuning
  • 4 Experimental Results
  • 4.1 Experimental Setups
  • 4.2 Results on Text Representation Tasks
  • 4.3 Results on Retrieval-Oriented Biomedical Applications
  • 4.4 Unsupervised Retrieval Performance
  • 4.5 Studies on Instruction Fine-tuning
  • 4.6 Additional Studies
  • 4.7 Case Study
  • 5 Conclusion
  • Acknowledgement
  • Limitation
  • Ethics Consideration
  • References
  • A Additional Synthetic Data Augmentation Details
  • A.1 Prompt format to Generate Query from Passage
  • A.2 Prompt Format to Generate Task and Pairs
  • A.3 Case Study
  • B Task and Dataset Information
  • B.1 Pre-training Corpus
  • B.2 Fine-tuning Task and Dataset
  • B.3 Evaluation Task and Dataset
  • C Baseline Information
  • C.1 Baselines for Retrieval Tasks in Main Experiments
  • C.2 Baselines for Retrieval-Oriented Downstream Applications
  • D Cosine Similarity v.s. Dot Product
  • E Similarity Score
  • F Efficiency

Knowls

  1. Knowl 1 — Two-stage BMRETRIEVER framework

    model/method

    BMRETRIEVER is a family of dense biomedical text retrievers built from autoregressive language models. The training procedure has two stages: (1) unsupervised contrastive pre-training on query–passage pairs constructed from large public biomedical corpora, followed by (2) supervised instruction fine-tuning on heterogeneous biomedical and general-domain retrieval data, augmented with synthetic examples generated by large language models. The resulting encoder is intended to generalize across standard information retrieval, sentence similarity, biomedical question answering, entity linking, and paper recommendation, including task formats that were not present during fine-tuning.

  2. Knowl 2 — Instruction-conditioned dense retrieval representation

    equation

    For a query qq and passage pp, let IqI_q and IpI_p be their task instructions, let ⊕\oplus denote string concatenation, and let EE be the BMRETRIEVER encoder. The query and passage embeddings are

    eq=E(Iq⊕q),ep=E(Ip⊕p).e_q = E(I_q \oplus q), \qquad e_p = E(I_p \oplus p).

    The relevance score is the dot product

    sim⁡(q,p)=eq⊤ep.\operatorname{sim}(q,p)=e_q^{\top}e_p.

    Because BMRETRIEVER uses autoregressive transformers, an end-of-sequence token is appended to each input, and the final-layer embedding of that token is used as eqe_q or epe_p. The model ranks passages by the resulting scalar dot-product scores; cosine similarity was evaluated but was not used by default because dot product performed better empirically.

  3. Knowl 3 — Unsupervised biomedical contrastive pre-training

    model/method

    BMRETRIEVER’s first training stage adapts the encoder to biomedical language using public unlabeled corpora containing biomedical publications, medical textbooks, clinical and biomedical resources, and general web text. For a corpus with document titles, the title is used as the query and its corresponding abstract as the positive passage. For an untitled corpus, two disjoint passages sampled from the same document are paired, with one treated as the query and the other as the positive passage. Passages belonging to other examples in the same mini-batch serve as in-batch negatives.

    For a mini-batch BB of positive pairs (qi,pi)(q_i,p_i), the pre-training objective is the InfoNCE loss

    Lcpt=−log⁡exp⁡(sim⁡(qi,pi)/τ)∑j∈Bexp⁡(sim⁡(qi,pj)/τ),\mathcal{L}_{\mathrm{cpt}}=-\log\frac{\exp(\operatorname{sim}(q_i,p_i)/\tau)}{\sum_{j\in B}\exp(\operatorname{sim}(q_i,p_j)/\tau)},

    where sim⁡\operatorname{sim} is the dot-product score, τ\tau is the temperature, and the experiments use τ=1\tau=1. The released pre-training recipe uses approximately 10 million passages or passage pairs; for the 7B model, only 1 million passages were used because of efficiency constraints.

  4. Knowl 4 — Instruction fine-tuning with labeled retrieval tasks

    model/method

    The second training stage combines biomedical and general-domain supervision at both sentence and passage levels. Biomedical sources include medical natural-language inference, medical question pairs, biomedical question answering, StackExchange answers, and medical dialogues. General-domain sources include MS MARCO, Natural Questions, FEVER, ELI5, GooAQ, and natural-language inference. Non-retrieval data are converted into retrieval pairs: questions are paired with gold evidence, entailment or similarity pairs are treated as positives, contradiction or dissimilarity pairs as hard negatives, and medical answers are treated as passages relevant to their user queries.

    For each training example, qiq_i is a query, pi+p_i^+ is its positive passage, and pi−p_i^- is a hard-negative passage. The fine-tuning loss for a mini-batch BB is an InfoNCE objective that uses both positive and hard-negative passages from every batch example:

    Lft=−log⁡exp⁡(sim⁡(qi,pi+)/τ)∑j∈B[exp⁡(sim⁡(qi,pj+)/τ)+exp⁡(sim⁡(qi,pj−)/τ)].\mathcal{L}_{\mathrm{ft}}=-\log\frac{\exp(\operatorname{sim}(q_i,p_i^+)/\tau)}{\sum_{j\in B}\left[\exp(\operatorname{sim}(q_i,p_j^+)/\tau)+\exp(\operatorname{sim}(q_i,p_j^-)/\tau)\right]}.

    Here τ=1\tau=1, and the denominator supplies both in-batch positive passages that are negative for qiq_i and explicitly mined hard negatives.

  5. Knowl 5 — LLM-generated synthetic retrieval augmentation

    algorithm

    BMRETRIEVER augments scarce biomedical supervision with two forms of synthetic data. GPT-3.5 is prompted with a biomedical passage from the pre-training corpus and generates a relevant query for that passage. GPT-4 is first prompted to brainstorm diverse biomedical retrieval scenarios and is then prompted to generate, for each scenario, a user query, a relevant passage, and a challenging irrelevant passage. The generated scenarios vary query length, ambiguity, topic, and required level of biomedical expertise.

    Input: Biomedical passages, GPT-3.5, GPT-4, and the E5-base retriever
    Output: Filtered positive and hard-negative retrieval pairs
    For each selected biomedical passage, prompt GPT-3.5 to generate a relevant query
    Generate approximately 500,000 query–passage pairs and retain about 420,000 after filtering
    Prompt GPT-4 to generate approximately 20,000 biomedical retrieval tasks
    For each generated task, prompt GPT-4 to produce a query, a positive passage, and a hard-negative passage
    For every labeled or synthetic query, retrieve the top 100 corpus passages with E5-base
    Randomly select one of these top-100 passages as a hard negative when a negative is needed
    For each synthetic positive pair, retain the query only if its positive passage is among E5-base's top three results
    Return the retained pairs for instruction fine-tuning
  6. Knowl 6 — Model scales and reproducible training configuration

    experimental setup

    BMRETRIEVER is released at four scales. The 410M and 1B models use Pythia backbones, the 2B model uses Gemma, and the 7B model uses BioMistral. Their configurations are:

    Model Parameters Backbone Layers Embedding dimension
    BMRETRIEVER-410M 410M Pythia 24 1024
    BMRETRIEVER-1B 1B Pythia 16 2048
    BMRETRIEVER-2B 2B Gemma 18 2048
    BMRETRIEVER-7B 7B BioMistral 32 4096

    Pre-training learning rates are 5×10−55\times10^{-5} for 410M and 1B, 4×10−54\times10^{-5} for 2B, and 2×10−52\times10^{-5} for 7B. Fine-tuning learning rates are 5×10−55\times10^{-5}, 5×10−55\times10^{-5}, 2×10−52\times10^{-5}, and 1×10−51\times10^{-5} at the same four scales. Global batch sizes are 256, 256, 128, and 64. Training uses AdamW, a 100-step linear warm-up, two pre-training epochs, one fine-tuning epoch, a maximum sequence length of 512 tokens, LoRA with rank r=16r=16 and scaling α=32\alpha=32, bfloat16 quantization, and DeepSpeed gradient checkpointing on four NVIDIA H100 GPUs.

    Evaluation covers eleven datasets and five task types: NFCorpus, SciFact, SciDocs, and TREC-COVID for standard information retrieval; BIOSSES for sentence similarity; BioASQ, PubMedQA, and iCliniq for question answering; DrugBank and MeSH for entity linking; and RELISH for paper recommendation. Metrics are nDCG@10 for standard retrieval, Spearman correlation for sentence similarity, Recall@5, Recall@20, and nDCG@20 for question answering, Recall@1, Recall@5, and MRR@5 for entity linking, and MAP and nDCG for paper recommendation. Training and test pairs do not overlap.

  7. Knowl 7 — Biomedical representation performance and parameter efficiency

    data/table

    The main representation evaluation compares standard biomedical retrieval on NFCorpus, SciFact, SciDocs, and TREC-COVID with biomedical sentence similarity on BIOSSES. ‘Avg. retrieval’ is the mean over the four standard retrieval datasets, and ‘Avg. all’ is the mean over all five tasks. The following values reproduce the BMRETRIEVER variants and representative competing models:

    Model Size NFCorpus SciFact SciDocs TREC-COVID BIOSSES Avg. retrieval / Avg. all
    E5-Large-v2 335M 0.371 0.726 0.201 0.665 0.836 0.560
    GTR-XXL 4.8B 0.342 0.662 0.161 0.501 0.819 0.497
    LLM2Vec 7B 0.393 0.788 0.225 0.776 0.852 0.606
    E5-Mistral 7B 0.386 0.764 0.162 0.872 0.855 0.608
    BMRETRIEVER-410M 410M 0.321 0.711 0.167 0.831 0.840 0.574
    BMRETRIEVER-1B 1B 0.344 0.760 0.180 0.840 0.858 0.596
    BMRETRIEVER-2B 2B 0.351 0.760 0.199 0.863 0.828 0.600
    BMRETRIEVER-7B 7B 0.364 0.778 0.201 0.861 0.847 0.610

    The BMRETRIEVER averages in the final column are respectively 0.574, 0.596, 0.600, and 0.610; the corresponding average-retrieval values are 0.321, 0.344, 0.351, and 0.364. Relative to BMRETRIEVER-7B, the 410M, 1B, and 2B variants retain 94.1%, 97.7%, and 98.4% of average-all performance while using 5.9%, 14.3%, and 28.6% as many parameters. BMRETRIEVER-410M exceeds the reported average-all scores of the 1B–5B baselines, including models with as many as 11.7 times more parameters, and BMRETRIEVER-2B reaches the performance range of models larger than 5B.

  8. Knowl 8 — Performance on retrieval-oriented biomedical applications

    data/table

    The downstream evaluation tests whether the embeddings transfer beyond the standard representation benchmarks. The columns are, in order: BioASQ Recall@5, Recall@20, nDCG@20; PubMedQA Recall@5, Recall@20, nDCG@20; iCliniq Recall@5, Recall@20, nDCG@20; DrugBank Recall@1, Recall@5, MRR@5; MeSH Recall@1, Recall@5, MRR@5; and RELISH MAP, nDCG. Values are percentages for recall and ranking metrics as reported by the study.

    Model BQ5 BQ20 BQN PQ5 PQ20 PQN IQ5 IQ20 IQN DB1 DB5 DBM MSH1 MSH5 MSHM RMAP RNDCG
    E5-Mistral 39.6 55.4 52.7 72.6 74.2 70.0 56.7 72.2 51.8 78.5 92.2 84.0 47.9 76.2 61.3 85.2 90.8
    BMRETRIEVER-410M 39.9 54.2 53.1 73.8 74.6 72.4 60.6 72.8 56.6 81.4 88.2 83.7 31.5 53.8 39.8 85.2 91.2
    BMRETRIEVER-1B 40.4 55.8 53.4 73.6 74.4 72.7 61.1 73.7 56.8 84.7 89.1 86.5 35.5 60.3 48.8 85.2 91.3
    BMRETRIEVER-2B 42.5 56.5 55.7 74.0 74.6 73.1 70.0 81.2 65.7 82.6 90.2 85.8 45.6 71.3 59.5 85.4 91.5
    BMRETRIEVER-7B 43.7 60.2 57.4 74.2 74.6 73.8 68.4 79.7 63.7 84.7 92.8 88.0 49.8 76.5 61.1 86.7 92.2

    BMRETRIEVER generally improves with scale and outperforms the strongest listed comparison on many question-answering, entity-linking, and paper-recommendation metrics. BMRETRIEVER-7B reaches 57.4 nDCG@20 on BioASQ, 73.8 on PubMedQA, 63.7 on iCliniq, 88.0 MRR@5 on DrugBank, 61.1 MRR@5 on MeSH, and 86.7 MAP on RELISH. The model also transfers to entity linking and paper recommendation even though those task formats were not included in instruction fine-tuning.

  9. Knowl 9 — Ablation evidence for pre-training, instructions, and data efficiency

    empirical result

    Removing any of BMRETRIEVER’s three principal components—instructions, biomedical contrastive pre-training, or supervised fine-tuning—reduces average performance across the five main representation tasks. Biomedical pre-training is especially helpful for smaller models, whereas larger backbones already contain more biomedical knowledge. Using cropping alone, meaning randomly pairing two passages from a document without title–abstract structure, performs worse than the full contrastive pre-training strategy.

    The unsupervised configuration, which uses unlabeled corpora for pre-training and synthetic data for fine-tuning but no human-labeled fine-tuning data, already outperforms most unsupervised and several supervised baselines:

    Model NFCorpus SciFact SciDocs TREC-COVID BIOSSES Avg. retrieval Avg. all
    Contriever 0.328 0.677 0.165 0.274 0.781 0.347 0.434
    COCO-DR 0.243 0.724 0.150 0.483 0.801 0.400 0.480
    E5-Large-v2 0.337 0.723 0.218 0.618 0.822 0.474 0.543
    LLM2Vec 0.271 0.687 0.153 0.557 0.832 0.417 0.500
    BMRETRIEVER-410M 0.306 0.677 0.180 0.802 0.834 0.491 0.560
    BMRETRIEVER-1B 0.330 0.744 0.187 0.800 0.833 0.515 0.579
    BMRETRIEVER-2B 0.342 0.738 0.198 0.848 0.847 0.531 0.593
    BMRETRIEVER-7B 0.355 0.750 0.208 0.833 0.861 0.537 0.601

    Reducing the training data to 10%, 50%, and 100% produces the following average-all scores. Pre-training values do not include later fine-tuning; fine-tuning values use the full pre-training checkpoint.

    Stage Model 10% 50% 100%
    Pre-training BMRETRIEVER-410M 0.540 0.554 0.560
    Pre-training BMRETRIEVER-1B 0.564 0.575 0.579
    Fine-tuning BMRETRIEVER-410M 0.562 0.571 0.574
    Fine-tuning BMRETRIEVER-1B 0.590 0.595 0.596

    Synthetic data produces the largest fine-tuning gain overall because it supplies more examples and broader task coverage. Biomedical fine-tuning is particularly useful for sentence similarity and dialogue-style tasks, while general-domain data mainly improves conventional short-query/long-passage retrieval.

  10. Knowl 10 — Operational and data-related limitations

    limitation

    Scaling BMRETRIEVER increases embedding-generation latency and storage requirements for passage vectors. The reported document encoding speeds and retrieval latencies are 471.2 documents per second per GPU and 14.6 ms for 410M, 194.0 and 28.6 ms for 1B, 166.2 and 28.6 ms for 2B, and 51.8 and 58.6 ms for 7B. These costs are not substantially worse than comparably sized baselines, but reducing inference latency and embedding storage remains future work.

    Synthetic supervision also requires paid GPT API calls; the total generation cost reported for BMRETRIEVER was below $500. Biomedical text generated by language models can introduce misinformation or hallucinations, although medical students found no such issue in a random evaluation of 200 generated examples. A string-matching check found no overlap between training and test queries, but the use of public biomedical corpora that may also occur in evaluation collections remains a potential contamination consideration.

Coverage note — Qualitative retrieval case studies, the full corpus and dataset inventories, cosine-similarity distribution plots, and the complete baseline descriptions were omitted because they corroborate the included method, quantitative results, and limitations without adding separate load-bearing contributions.

References

  1. 1.Chris Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, and Michael Collins. 2019. Synthetic QA corpora generation with roundtrip consistency. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6168–6173, Florence, Italy. Association for Computational Linguistics.
  2. 2.Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2023. Task-aware retrieval with instructions. In Findings of the Association for Computational Linguistics: ACL 2023, pages 3650–3675, Toronto, Canada. Association for Computational Linguistics.
  3. 3.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. MS MARCO: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268.
  4. 4.Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2vec: Large language models are secretly powerful text encoders. In First Conference on Language Modeling.
  5. 5.Asma Ben Abacha, Chaitanya Shivade, and Dina Demner-Fushman. 2019. Overview of the MEDIQA 2019 shared task on textual inference, question entailment and question answering. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 370–379, Florence, Italy. Association for Computational Linguistics.
  6. 6.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 2397–2430. PMLR.
  7. 7.Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. In European Conference on Information Retrieval, pages 716–722. Springer.
  8. 8.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  9. 9.Peter Brown, Aik-Choon Tan, Mohamed A El-Esawi, Thomas Liehr, Oliver Blanck, Douglas P Gladue, Gabriel MF Almeida, Tomislav Cernava, Carlos O Sorzano, Andy WK Yeung, et al. 2019. Large expert-curated database for benchmarking document similarity detection in biomedical literature search. Database, 2019.
  10. 10.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  11. 11.Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216.
  12. 12.Qingyu Chen, Alexis Allot, and Zhiyong Lu. 2021. Litcovid: an open database of covid-19 literature. Nucleic acids research, 49(D1):D1534–D1540.
  13. 13.Shu Chen, Zeqian Ju, Xiangyu Dong, Hongchao Fang, Sicheng Wang, Yue Yang, Jiaqi Zeng, Ruisi Zhang, Ruoyu Zhang, Meng Zhou, Penghui Zhu, and Pengtao Xie. 2020. Meddialog: A large-scale medical dialogue dataset. CoRR, abs/2004.03329.
  14. 14.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020. SPECTER: Document-level representation learning using citation-informed transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270–2282, Online. Association for Computational Linguistics.
  15. 15.Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot dense retrieval from 8 examples. In The Eleventh International Conference on Learning Representations.
  16. 16.Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990. Indexing by latent semantic analysis. Journal of the American society for information science, 41(6):391–407.
  17. 17.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558–3567, Florence, Italy. Association for Computational Linguistics.
  18. 18.Nicolas Fiorini, Robert Leaman, David J Lipman, and Zhiyong Lu. 2018. How user intelligence is improving pubmed. Nature biotechnology, 36(10):937–945.
  19. 19.Giacomo Frisoni, Miki Mizutani, Gianluca Moro, and Lorenzo Valgimigli. 2022. BioReader: a retrieval-enhanced text-to-text transformer for biomedical literature. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5770–5793, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  20. 20.Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, and Diego Garcia-Olano. 2019. Learning dense representations for entity retrieval. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 528–537, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  22. 22.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  23. 23.Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry P. Heck. 2013. Learning deep structured semantic models for web search using clickthrough data. In 22nd ACM International Conference on Information and Knowledge Management, CIKM’13, San Francisco, CA, USA, October 27 - November 1, 2013, pages 2333–2338. ACM.
  24. 24.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  25. 25.Fan Jiang, Tom Drummond, and Trevor Cohn. 2023. Noisy self-training with synthetic queries for dense retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11991–12008, Singapore. Association for Computational Linguistics.
  26. 26.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  27. 27.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. PubMedQA: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2567–2577, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11):btad651.
  29. 29.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  30. 30.Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch. 2021. GooAQ: Open question answering with diverse answer types. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 421–433, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  32. 32.Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. BioMistral: A collection of open-source pretrained large language models for medical domains. In Findings of the Association for Computational Linguistics ACL 2024, pages 5848–5864, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  33. 33.Jakub Lála, Odhran O’Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White. 2023. Paperqa: Retrieval-augmented generative agent for scientific research. arXiv preprint arXiv:2312.07559.
  34. 34.Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. Nv-embed: Improved techniques for training llms as generalist embedding models. Preprint, arXiv:2405.17428.
  35. 35.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  36. 36.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc.
  37. 37.Chaofan Li, Zheng Liu, Shitao Xiao, and Yingxia Shao. 2023a. Making large language models a better foundation for dense retrieval. arXiv preprint arXiv:2312.15503.
  38. 38.Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023b. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6).
  39. 39.Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. How to train your dragon: Diverse augmentation towards generalizable dense retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6385–6400, Singapore. Association for Computational Linguistics.
  40. 40.Carolyn E Lipscomb. 2000. Medical subject headings (mesh). Bulletin of the Medical Library Association, 88(3):265.
  41. 41.Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Collier. 2021. Self-alignment pretraining for biomedical entity representations. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4228–4238, Online. Association for Computational Linguistics.
  42. 42.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969–4983, Online. Association for Computational Linguistics.
  43. 43.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  44. 44.Man Luo, Arindam Mitra, Tejas Gokhale, and Chitta Baral. 2022a. Improving biomedical information retrieval with neural retrievers. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11038–11046.
  45. 45.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022b. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6):bbac409.
  46. 46.Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2021. Zero-shot neural passage retrieval via domain-targeted synthetic question generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1075–1088, Online. Association for Computational Linguistics.
  47. 47.Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. Fine-tuning llama for multi-stage text retrieval. arXiv preprint arXiv:2310.08319.
  48. 48.Clara H McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain. 2020. Effective transfer learning for identifying similar questions: matching user questions to covid-19 faqs. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, pages 3458–3465.
  49. 49.Rui Meng, Ye Liu, Semih Yavuz, Divyansh Agarwal, Lifu Tu, Ning Yu, Jianguo Zhang, Meghana Bhat, and Yingbo Zhou. 2022. Augtriever: Unsupervised dense retrieval by scalable data augmentation. arXiv preprint arXiv:2212.08841.
  50. 50.Sunil Mohan, Nicolas Fiorini, Sun Kim, and Zhiyong Lu. 2017. Deep learning for biomedical information retrieval: Learning textual relevance from click logs. In BioNLP 2017, pages 222–231, Vancouver, Canada,. Association for Computational Linguistics.
  51. 51.Niklas Muennighoff. 2022. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904.
  52. 52.Aakanksha Naik, Sravanthi Parasa, Sergey Feldman, Lucy Lu Wang, and Tom Hope. 2022. Literature-augmented clinical outcome prediction. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 438–453, Seattle, United States. Association for Computational Linguistics.
  53. 53.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005.
  54. 54.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844–9855, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  55. 55.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2023. Med-HALT: Medical domain hallucination test for large language models. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 314–334, Singapore. Association for Computational Linguistics.
  56. 56.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5835–5847, Online. Association for Computational Linguistics.
  57. 57.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67.
  58. 58.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506.
  59. 59.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  60. 60.Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023. NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10776–10787, Singapore. Association for Computational Linguistics.
  61. 61.Wenqi Shi, Yuchen Zhuang, Yuanda Zhu, Henry Iwinski, Michael Wattenbarger, and May Dongmei Wang. 2023. Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making. In Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, pages 1–10.
  62. 62.Chaitanya Shivade. 2017. Mednli — a natural language inference dataset for the clinical domain.
  63. 63.Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. 2023. SciRepEval: A multi-format benchmark for scientific document representations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5548–5566, Singapore. Association for Computational Linguistics.
  64. 64.Gizem Sogancıoğlu, Hakime Öztürk, and Arzucan Özgür. 2017. Biosses: a semantic sentence similarity estimation system for the biomedical domain. Bioinformatics, 33(14):i49–i58.
  65. 65.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. One embedder, any task: Instruction-finetuned text embeddings. In Findings of the Association for Computational Linguistics: ACL 2023, pages 1102–1121, Toronto, Canada. Association for Computational Linguistics.
  66. 66.Flax Sentence Embeddings Team. 2021. Stack exchange question pairs.
  67. 67.Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295.
  68. 68.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  69. 69.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  70. 70.George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015. An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16(1):1–28.
  71. 71.Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: Constructing a pandemic information retrieval test collection. SIGIR Forum, 54(1).
  72. 72.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. Association for Computational Linguistics.
  73. 73.Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych. 2022a. GPL: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2345–2360, Seattle, United States. Association for Computational Linguistics.
  74. 74.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022b. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533.
  75. 75.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Improving text embeddings with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11897–11916, Bangkok, Thailand. Association for Computational Linguistics.
  76. 76.Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Nancy Xin Ru Wang, Christopher Wilhelm, Boya Xie, Douglas M. Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020. CORD-19: The COVID-19 open research dataset. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020, Online. Association for Computational Linguistics.
  77. 77.Yubo Wang, Xueguang Ma, and Wenhu Chen. 2023. Augmenting black-box llms with medical textbooks for clinical question answering. arXiv preprint arXiv:2309.02233.
  78. 78.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4003–4012, Marseille, France. European Language Resources Association.
  79. 79.David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al. 2018. Drugbank 5.0: a major update to the drugbank database for 2018. Nucleic acids research, 46(D1):D1074–D1082.
  80. 80.Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024. Benchmarking retrieval-augmented generation for medicine. In Findings of the Association for Computational Linguistics ACL 2024, pages 6233–6251, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  81. 81.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In International Conference on Learning Representations.
  82. 82.Ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang, Bowen Jin, May Dongmei Wang, Joyce Ho, and Carl Yang. 2024. RAM-EHR: Retrieval augmentation meets clinical predictions on electronic health records. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 754–765, Bangkok, Thailand. Association for Computational Linguistics.
  83. 83.Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. RankRAG: Unifying context ranking with retrieval-augmented generation in LLMs. In Thirty-eighth Annual Conference on Neural Information Processing Systems.
  84. 84.Yue Yu, Chenyan Xiong, Si Sun, Chao Zhang, and Arnold Overwijk. 2022. COCO-DR: Combating distribution shift in zero-shot dense retrieval with contrastive and distributionally robust learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1462–1479, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  85. 85.Tianjun Zhang, Shishir G Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E. Gonzalez. 2024. RAFT: Adapting language model to domain specific RAG. In First Conference on Language Modeling.
  86. 86.Yu Zhang, Hao Cheng, Zhihong Shen, Xiaodong Liu, Ye-Yi Wang, and Jianfeng Gao. 2023. Pre-training multi-task contrastive learning models for scientific literature understanding. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12259–12275, Singapore. Association for Computational Linguistics.

Citation

MLA
Xu, R., et al. “BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 22234–54, https://doi.org/10.18653/v1/2024.emnlp-main.1241.
APA
Xu, R., Shi, W., Yu, Y., Zhuang, Y., Zhu, Y., Wang, M. D., Ho, J. C., Zhang, C., & Yang, C. (2024). BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 22234–22254. https://doi.org/10.18653/v1/2024.emnlp-main.1241
Chicago
Xu, R., W. Shi, Y. Yu, et al. 2024. “BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 22234–54. https://doi.org/10.18653/v1/2024.emnlp-main.1241.
Harvard
Xu, R. et al. (2024) “BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 22234–22254. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1241.
Vancouver
1. Xu R, Shi W, Yu Y, Zhuang Y, Zhu Y, Wang MD, Ho JC, Zhang C, Yang C (2024) BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 22234–22254

BibTeX

@inproceedings{xu-etal-2024-bmretriever,
    title = "{BMR}etriever: Tuning Large Language Models as Better Biomedical Text Retrievers",
    author = "Xu, Ran  and
      Shi, Wenqi  and
      Yu, Yue  and
      Zhuang, Yuchen  and
      Zhu, Yanqiao  and
      Wang, May Dongmei  and
      Ho, Joyce C.  and
      Zhang, Chao  and
      Yang, Carl",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1241/",
    doi = "10.18653/v1/2024.emnlp-main.1241",
    pages = "22234--22254"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/