Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Jiwoong SohnYein ParkChanwoong YoonSihyeon ParkHyeon HwangMujeen SungHyunjae KimJaewoo Kang

article2025NAACL50 citations

Proposes a biomedical question-answering framework that boosts accuracy by using model-generated rationales for query formulation, balancing retrieval across diverse medical corpora, and filtering out distracting context with a perplexity-trained lightweight model.

Listen

Large language models show great promise for medical applications, but their deployment in high-stakes clinical settings is hindered by factual errors, called hallucinations, and outdated parametric knowledge. While retrieval-augmented generation connects models to external medical knowledge, conventional frameworks often underperform in medicine because specialized queries are hard to formulate, dense retrievers over-index on large corpora, and models are easily misled by irrelevant or unhelpful reference snippets.

The article evaluates a novel retrieval-augmented framework, RAG2, designed to enhance the reliability and accuracy of medical question answering. The objective is to demonstrate that combining rationale-based querying, balanced multi-corpus retrieval, and confidence-based document filtering improves language model performance across varied model architectures and sizes without requiring costly model retraining.

To evaluate this framework, the authors conducted empirical benchmarks across three major multiple-choice medical examination datasets comprising over 200,000 total questions: MedQA, MedMCQA, and MMLU-Med, as well as a real-world set of open-ended clinical queries. They tested open-source models, medically specialized models, and leading commercial models. The approach formulates queries using step-by-step reasoning generated by the model, extracts snippets equally across four diverse biomedical corpora, and employs a lightweight filtering model trained on model uncertainty signals, specifically perplexity reductions, to prune unhelpful context before final single-pass generation.

The findings establish that the proposed framework delivers consistent, notable gains across benchmarks. First, the framework improved the average accuracy of the open-source baseline by 6.1 percentage points, the specialized medical model by 3.8 percentage points, and the leading commercial model by 0.9 percentage points. Second, it outperformed existing state-of-the-art medical retrieval frameworks by up to 5.6 percentage points on the open-source model. Third, the small filtering model effectively matched the filtering accuracy of a large commercial system while eliminating recurring application programming interface costs and expensive iterative generation cycles. Fourth, balanced multi-source retrieval consistently surpassed single-corpus and stacked retrieval approaches by preventing dominant corpora from overshadowing critical guidelines and textbooks.

These results demonstrate that simply retrieving more medical text can actively degrade model accuracy if distractor content is not rigorously filtered out. For healthcare organizations and technology leaders, the proposed architecture provides a computationally efficient path to improve diagnostic accuracy, reduce misdiagnosis risks, and control operational serving costs. High-quality single-pass filtering proves to be a safer, lower-latency alternative to multi-step recursive reasoning architectures.

Organizations developing medical AI systems should adopt structured multi-corpus balancing and integrate lightweight filtering modules based on model confidence signals rather than relying on standard similarity-based search. Before operational deployment in clinical workflows, stakeholders should pilot these pipelines on broader specialized tasks and establish validation guardrails for cases where the initial reasoning steps are flawed.

Confidence in these findings is high for multiple-choice medical examinations, but readers should note key limitations. The framework was evaluated primarily on multiple-choice formats within the biomedical domain, tested only one compact filtering model size, and evaluated snippets individually rather than jointly. Further validation on complex, real-world conversational workflows across additional domains remains necessary.

arXiv: 2411.00300dmis-lab/RAG2
Cover for Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Abstract

Large language models (LLM) hold significant potential for applications in biomedicine, but they struggle with hallucinations and outdated knowledge. While retrieval-augmented generation (RAG) is generally employed to address these issues, it also has its own set of challenges: (1) LLMs are vulnerable to irrelevant or unhelpful context, (2) medical queries are often not well-targeted for helpful information, and (3) retrievers are prone to bias toward the specific source corpus they were trained on. In this study, we present RAG² (Rationale-Guided RAG), a new framework for enhancing the reliability of RAG in biomedical contexts. RAG² incorporates three key innovations: a small filtering model trained on perplexity-based labels of rationales, which selectively augments informative snippets of documents while filtering out distractors; LLM-generated rationales as queries to improve the utility of retrieved snippets; a structure designed to retrieve snippets evenly from a comprehensive set of four biomedical corpora, effectively mitigating retriever bias. Our experiments demonstrate that RAG² improves the state-of-the-art LLMs of varying sizes, with improvements of up to 6.1%, and it outperforms the previous best medical RAG model by up to 5.6% across three medical question-answering benchmarks. Our code is available at https://github.com/dmis-lab/RAG2

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Retrieval-Augmented Generation
  • 2.2 Medical RAG
  • 3 Method
  • 3.1 Objective
  • 3.2 Rationale-Guided Filtering
  • 3.3 Rationale-Based Query Formulation
  • 3.4 Balanced Retrieval
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Models
  • 4.3 Main Results
  • 5 Analysis
  • 5.1 Ablation Study
  • 5.2 Quality of Rationale Queries
  • 5.3 Case Study
  • 6 Conclusion
  • Limitation
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Ablation Study of Top-k documents on MMLU-Med
  • A.2 Datasets
  • A.3 Implementation Details
  • A.4 Open-ended Clinical Questions
  • A.4.1 Metrics
  • A.4.2 Results
  • A.5 Balanced Retrieval

Knowls

  1. Knowl 1 — RAG2 Architecture for Medical Question Answering

    model/method

    RAG2\text{RAG}^2 (Rationale-Guided Retrieval-Augmented Generation) is a framework designed for medical question answering (QA) that mitigates retriever corpus bias, distractors from unhelpful context, and poorly targeted queries without requiring fine-tuning of the underlying generator large language model (LLM). The framework consists of three sequential modules:

    1. Rationale-Based Query Formulation: Given an input medical question xx, the base LLM produces a step-by-step reasoning rationale via chain-of-thought prompting. The generated rationale text alone serves as the search query for the retriever rather than the original question.
    2. Balanced Multi-Corpus Retrieval and Reranking: Documents are retrieved in equal quotas across four separate biomedical corpora (PubMed abstracts, PubMed Central full texts, Clinical Practice Guidelines, and medical textbooks) using the MedCPT dense retriever to prevent large corpora from overshadowing smaller, specialized resources. Candidate snippets are subsequently reranked using a MedCPT cross-encoder scoring the original query against each snippet.
    3. Rationale-Guided Filtering: A compact filtering model (Flan-T5-large, 770M parameters), trained on perplexity-differential labels derived from base LLM rationale generations, evaluates candidate snippets individually to discard distractors and retain only evidence that improves model confidence and accuracy.

    The filtered snippets and the initial question are concatenated and passed to the base LLM in a single-pass inference step to generate the final response.

  2. Knowl 2 — Perplexity-Differential Labeling for Document Filtering

    equation

    To evaluate whether a candidate document snippet dd assists a target large language model on a medical input query xx, the RAG2\text{RAG}^2 filtering model is trained on labels determined by the change in perplexity ΔPPL\Delta\text{PPL} of the model-generated rationale:

    ΔPPL=PPL(x)−PPL(x,d)\Delta\text{PPL} = \text{PPL}(x) - \text{PPL}(x, d)

    where sequence perplexity for an LL-token rationale sequence x=(x0,x1,…,xL−1)x = (x_0, x_1, \dots, x_{L-1}) under language model conditional probability distribution PP is computed as:

    PPL(x)=exp⁡(−1L∑i=0L−1log⁡P(xi∣x<i))\text{PPL}(x) = \exp\left(-\frac{1}{L} \sum_{i=0}^{L-1} \log P(x_i \mid x_{<i})\right)

    PPL(x,d)=exp⁡(−1L∑i=0L−1log⁡P(xi∣x<i,d))\text{PPL}(x, d) = \exp\left(-\frac{1}{L} \sum_{i=0}^{L-1} \log P(x_i \mid x_{<i}, d)\right)

    Training labels for the filtering classifier are assigned using model correctness and a threshold τ\tau set to select the top 25% of perplexity differentials:

    • If the LLM answers correctly without retrieval: the snippet is labeled Helpful if the LLM remains correct with retrieval and ΔPPL≥τ\Delta\text{PPL} \ge \tau; it is Discarded if correct with retrieval but ΔPPL<τ\Delta\text{PPL} < \tau; and labeled Not Helpful if retrieval causes an incorrect answer.
    • If the LLM answers incorrectly without retrieval: the snippet is labeled Helpful if retrieval leads to a correct answer; Discarded if it remains incorrect but ΔPPL≥τ\Delta\text{PPL} \ge \tau; and labeled Not Helpful if it remains incorrect and ΔPPL<τ\Delta\text{PPL} < \tau.
  3. Knowl 3 — Rationale-Based Query Formulation via Chain-of-Thought Prompting

    model/method

    Medical examination questions often present either excessively long clinical vignettes (detailed patient history, physical examination, vitals, lab results) or overly brief prompts lacking explicit context. RAG2\text{RAG}^2 replaces the raw question xx with an LLM-generated rationale as the query for document retrieval.

    The base LLM generates the rationale using the following chain-of-thought prompt:

    "The following are multiple choice questions about medical knowledge. Solve them in a step-by-step fashion, starting by summarizing the available information. Output your explanation and single option from the given options as the final answer. Here is the question: [initial_query]"

    The extracted explanation text (excluding the initial query) is passed directly to the dense retriever. By isolating the rationale, the query expands brief questions with medical concepts and condenses verbose patient descriptions into key diagnostic hypotheses while staying within the maximum token capacity of the dense retriever encoder.

  4. Knowl 4 — Balanced Multi-Corpus Retrieval and Cross-Encoder Reranking

    model/method

    In biomedical retrieval, indexing multiple corpora into a single combined collection causes dense retrievers to disproportionately favor massive corpora (e.g., PubMed) over compact, clinically authoritative sources (e.g., medical textbooks and clinical guidelines). RAG2\text{RAG}^2 resolves this retriever bias by implementing balanced retrieval across four biomedical corpora:

    1. PubMed Abstracts: 36.5 million documents segmented into 69.7 million passages (400 GB index size).
    2. PubMed Central (PMC) Full Texts: 1.1 million full-text papers segmented via sliding window into 46.3 million passages (160 GB index size).
    3. Clinician Practical Guidelines (CPG): 35.7 thousand clinical practice guidelines segmented into 607.0 thousand passages (3.5 GB index size).
    4. Medical Textbooks: 18 medical textbooks segmented into 134.0 thousand passages (0.7 GB index size).

    An equal number of candidate passages (k/4k/4) is retrieved from each of the four corpora using the MedCPT dense retriever. The concatenated candidate set of kk snippets is then reranked with the MedCPT cross-encoder using the original question and each snippet to prioritize the most relevant evidence prior to filtering.

  5. Knowl 5 — Benchmark Accuracy Comparison of RAG2 across Medical QA Datasets

    data/table

    The RAG2\text{RAG}^2 framework was evaluated on three closed-book multiple-choice medical QA benchmarks: MedQA (1,273 USMLE test questions), MedMCQA (6,150 Indian medical entrance exam test questions), and MMLU-Med (1,089 questions across 6 biomedical subjects). It was compared against base LLMs without retrieval and with six retrieval/RAG baselines: MedCPT (k=1k=1), MedCPT+Rationale (k=1k=1), MedRAG (hybrid dense/sparse retrieval + reciprocal rank fusion), query2doc (k=1k=1), Adaptive-RAG (k=1k=1), and InstructRAG-ICL (2-shot, k=5k=5).

    Model MedQA MedMCQA MMLU-Med Average
    Open-source LLMs (0-shot)
    Llama-3-8B-Instruct 57.7 53.5 69.5 60.2
    + MedCPT (k=1k=1) 55.3 51.3 65.8 57.5
    + MedCPT+Rationale (k=1k=1) 58.0 52.1 70.3 60.1
    + MedRAG 56.4 56.6 69.2 60.7
    + query2doc (k=1k=1) 54.3 50.0 58.5 54.3
    + Adaptive-RAG 57.3 53.1 70.3 60.2
    + InstructRAG-ICL 55.5 55.7 71.9 61.8
    + RAG2\text{RAG}^2 (Ours) 64.6 59.4 74.8 66.3
    Medical LLMs (fine-tuned)
    Meerkat-7B 71.2 60.8 73.8 68.6
    + MedCPT (k=1k=1) 71.8 57.9 74.0 67.9
    + MedCPT+Rationale (k=1k=1) 73.3 58.4 75.7 69.1
    + MedRAG 67.9 60.6 76.1 68.2
    + query2doc (k=1k=1) 70.3 53.8 73.6 65.9
    + Adaptive-RAG 71.4 60.5 74.0 68.6
    + InstructRAG-ICL 65.8 53.2 63.7 60.9
    + RAG2\text{RAG}^2 (Ours) 75.6 63.0 78.7 72.4
    Commercial LLMs (0-shot)
    GPT-4o 88.5 76.7 92.8 86.0
    + MedCPT (k=1k=1) 86.6 72.5 90.1 83.1
    + MedCPT+Rationale (k=1k=1) 87.3 74.7 90.2 84.1
    + MedRAG 88.3 75.9 92.4 85.5
    + query2doc (k=1k=1) 89.1 73.9 91.5 84.8
    + Adaptive-RAG 88.5 76.7 92.5 85.9
    + InstructRAG-ICL 87.7 73.5 90.0 85.6
    + RAG2\text{RAG}^2 (Ours) 91.1 77.2 92.5 86.9

    RAG2\text{RAG}^2 achieved the highest accuracy across all model sizes, with average gains of +6.1% on Llama-3-8B-Instruct, +3.8% on Meerkat-7B, and +0.9% on GPT-4o over baseline models without RAG, and outperforming MedRAG by up to +5.6% on Llama-3-8B-Instruct. On MMLU-Med (an out-of-domain evaluation where the filter was trained on MedMCQA), RAG2\text{RAG}^2 achieved gains of +5.3% on Llama-3-8B-Instruct and +4.9% on Meerkat-7B.

  6. Knowl 6 — Comparison of Document Filtering Methods for RAG

    data/table

    To isolate the performance contribution of filtering strategies, Llama-3-8B-Instruct was evaluated on MedQA and MedMCQA with top-1 retrieved documents from MedCPT using the original query (InstructRAG-ICL used top-5 in-context demonstrations).

    Filtering Method MedQA (%) MedMCQA (%)
    Llama-3-8B (MedCPT w/o Filtering) 55.3 51.3
    + Adaptive-RAG 57.3 53.1
    + InstructRAG-ICL 55.5 55.7
    + GPT-4o 58.5 55.8
    + Rationale-guided Filtering (Flan-T5-large) 58.6 55.8

    Rationale-guided filtering using a 770M-parameter Flan-T5-large model trained on perplexity differentials outperformed Adaptive-RAG (which relies on coarse correct/incorrect response labels) by +1.3% on MedQA and +2.7% on MedMCQA. It achieved parity with GPT-4o filtering without requiring proprietary LLM API calls during inference.

  7. Knowl 7 — Impact of Rationale Query Quality on Medical Retrieval and Generation

    data/table

    The quality of the rationale used as the retrieval query affects downstream QA accuracy. On the MedQA test set, rationale queries produced by stronger generator LLMs consistently improved performance across different backbone generators:

    Generator Model Rationale Query Generated By
    Llama-3-8B Meerkat-7B GPT-4o
    Llama-3-8B-Instruct 63.4 71.5 73.6
    Meerkat-7B 71.3 74.6 78.8
    GPT-4o 87.4 88.3 89.8

    Rationales generated by GPT-4o provided the highest performance improvements for all backbone models, followed by Meerkat-7B and Llama-3-8B-Instruct, demonstrating that higher reasoning fidelity in the generated rationale translates into more targeted retrieval queries.

  8. Knowl 8 — Performance Comparison of Multi-Corpus Retrieval Strategies

    empirical result

    Comparison of retrieval configurations using top-1 retrieved documents on Llama-3-8B-Instruct and Meerkat-7B demonstrates the necessity of balanced multi-corpus retrieval over single-corpus and stacked multi-corpus indexing:

    • Single-Corpus Retrieval: Retrieving independently from a single corpus (PubMed, PMC, CPG clinical guidelines, or medical textbooks alone) yields inferior accuracy (between 53.1% and 60.8% for Llama-3-8B on MedQA), indicating that no individual corpus possesses comprehensive medical coverage.
    • Stacked Multi-Corpus Retrieval (MedRAG): Combining all corpora into a unified index yields 51.9% on MedQA, 49.0% on MedMCQA, and 65.1% on MMLU-Med for Llama-3-8B-Instruct, and 63.2%, 56.0%, and 73.1% respectively for Meerkat-7B.
    • Balanced Multi-Corpus Retrieval: Retrieving an equal quota from each of the four corpora followed by cross-encoder reranking achieves 55.3% on MedQA (+3.4% over MedRAG), 51.3% on MedMCQA (+2.3%), and 65.8% on MMLU-Med (+0.7%) for Llama-3-8B-Instruct. For Meerkat-7B, it achieves 71.8% on MedQA (+8.6% over MedRAG), 57.9% on MedMCQA (+1.9%), and 74.0% on MMLU-Med (+0.9%).
    • Balanced Retrieval with Rationale-Guided Filtering: Incorporating Flan-T5 filtering further increases accuracy to 64.6% on MedQA, 59.4% on MedMCQA, and 74.8% on MMLU-Med for Llama-3-8B-Instruct, and 75.6%, 63.0%, and 78.7% respectively for Meerkat-7B.
  9. Knowl 9 — Generalization of RAG2 to Open-Ended Clinical Queries

    empirical result

    When evaluated on the ClinicalQA25 benchmark—a dataset comprising 25 complex open-ended medical queries written by clinicians—using Llama-3-8B-Instruct as the backbone generator:

    • No RAG: Achieved a ROUGE-L score of ∼0.165\sim 0.165 and a BERTScore of ∼0.575\sim 0.575.
    • MedRAG: Achieved ROUGE-L of ∼0.185\sim 0.185 and BERTScore of ∼0.590\sim 0.590.
    • Balanced Retrieval (No Filter): Achieved ROUGE-L of ∼0.190\sim 0.190 and BERTScore of ∼0.600\sim 0.600.
    • RAG2\text{RAG}^2 (Full Pipeline): Achieved the highest performance with ROUGE-L of ∼0.205\sim 0.205 and BERTScore of ∼0.605\sim 0.605.

    These results demonstrate that the RAG2\text{RAG}^2 framework and its compact filtering model, despite being trained entirely on multiple-choice question-answering pairs, generalize effectively to open-ended, real-world clinical text generation.

  10. Knowl 10 — Limitations of the RAG2 Framework

    limitation

    The RAG2\text{RAG}^2 framework has several acknowledged limitations:

    1. Domain Generality: The methodology has been validated exclusively on biomedical and clinical benchmarks; its effectiveness on general-domain open-domain question answering has not been evaluated.
    2. Filtering Architecture Diversity: Experiments only utilized one model family and size (Flan-T5-large, 770M parameters) as the filter; alternative filter architectures and scales remain unexplored.
    3. Independent Document Evaluation: Perplexity-differential labeling and inference-stage filtering evaluate candidate snippets individually rather than jointly due to the filter's context length limitations, which may fail to capture inter-document interactions, context synergies, or multi-document redundancies.
    4. Sensitivity to Flawed Rationales: If the base LLM produces an erroneous initial rationale, the dense retriever may be guided toward irrelevant or misleading evidence snippets.

Coverage note — All primary contributed methods, formulations, benchmark tables, ablation studies, and stated limitations are included as dedicated knowls; qualitative case study details (COPD case example) were omitted as their principle is fully captured in the filtering method knowls.

References

  1. 1.AI@Meta. 2024. Llama 3 model card.
  2. 2.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations.
  3. 3.Anthony Chen, Pallavi Gudipati, Shayne Longpre, Xiao Ling, and Sameer Singh. 2021. Evaluating entity disambiguation and the role of popularity in retrieval-based nlp. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4472–4485.
  4. 4.Zeming Chen, Alejandro Hernández Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas Köpf, Amirkeivan Mohtashami, et al. 2023. Meditron-70b: Scaling medical pretraining for large language models. arXiv preprint arXiv:2311.16079.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53.
  6. 6.Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pages 758–759.
  7. 7.Sunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu, Xiaolin Hu, Yong Liu, Xiao Zhang, Gang Wang, and Jun Xu. 2024. Neural retrievers are biased towards llm-generated content. In ICLR 2024 Workshop: How Far Are We From AGI.
  8. 8.Chunjing Gan, Dan Yang, Binbin Hu, Hanxiao Zhang, Siyuan Li, Ziqi Liu, Yue Shen, Lin Ju, Zhiqiang Zhang, Jinjie Gu, et al. 2024. Similarity is not all you need: Endowing retrieval augmented generation with multi layered thoughts. arXiv preprint arXiv:2405.19893.
  9. 9.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
  10. 10.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  11. 11.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  12. 12.Rolf Jagerman, Honglei Zhuang, Zhen Qin, Xuanhui Wang, and Michael Bendersky. 2023. Query expansion by prompting large language models. arXiv preprint arXiv:2305.03653.
  13. 13.Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, and Jaewoo Kang. 2024a. Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models. Bioinformatics.
  14. 14.Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024b. Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7029–7043.
  15. 15.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  16. 16.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023b. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7969–7992.
  17. 17.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  18. 18.Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11):btad651.
  19. 19.Minki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi, and Sung Ju Hwang. 2024. Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks. Advances in Neural Information Processing Systems, 36.
  20. 20.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, Kentaro Inui, et al. 2024. Realtime qa: what’s the answer right now? Advances in Neural Information Processing Systems, 36.
  21. 21.Hyunjae Kim, Hyeon Hwang, Jiwoo Lee, Sihyeon Park, Dain Kim, Taewhoo Lee, Chanwoong Yoon, Jiwoong Sohn, Donghee Choi, and Jaewoo Kang. 2024. Small language models learn enhanced reasoning skills from medical textbooks. arXiv preprint arXiv:2404.00376.
  22. 22.Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon, and Jaewoo Kang. 2023. Automatic creation of named entity recognition datasets by querying phrase representations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7148–7163.
  23. 23.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611–626.
  24. 24.Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. BioMistral: A collection of open-source pretrained large language models for medical domains. In Findings of the Association for Computational Linguistics ACL 2024, pages 5848–5864, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  25. 25.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  26. 26.Cui Long, Yongbin Liu, Chunping Ouyang, and Ying Yu. 2024. Bailicai: A domain-optimized retrieval-augmented generation framework for medical applications. arXiv preprint arXiv:2407.21055.
  27. 27.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919.
  28. 28.OpenAI. 2022. Introducing chatgpt.
  29. 29.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 1.
  30. 30.OpenAI. 2024. Hello gpt-4o.
  31. 31.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248–260. PMLR.
  32. 32.Lin CY Rouge. 2004. A package for automatic evaluation of summaries. In Proceedings of Workshop on Text Summarization of ACL, Spain, volume 5.
  33. 33.Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, et al. 2024. Capabilities of gemini models in medicine. arXiv preprint arXiv:2404.18416.
  34. 34.Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024. Ares: An automated evaluation framework for retrieval-augmented generation systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 338–354.
  35. 35.Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature, 620(7972):172–180.
  36. 36.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  37. 37.Liang Wang, Nan Yang, and Furu Wei. 2023a. Query2doc: Query expansion with large language models. In The 2023 Conference on Empirical Methods in Natural Language Processing.
  38. 38.Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023b. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377.
  39. 39.Zihao Wang, Anji Liu, Haowei Lin, Jiaqi Li, Xiaojian Ma, and Yitao Liang. 2024a. Rat: Retrieval augmented thoughts elicit context-aware reasoning in long-horizon generation. arXiv preprint arXiv:2403.05313.
  40. 40.Zilong Wang, Zifeng Wang, Long Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang, Anush Mattapalli, Ankur Taly, Jingbo Shang, et al. 2024b. Speculative rag: Enhancing retrieval augmented generation through drafting. arXiv preprint arXiv:2407.08223.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  42. 42.Zhepei Wei, Wei-Lin Chen, and Yu Meng. 2024. Instructrag: Instructing retrieval-augmented generation with explicit denoising. arXiv preprint arXiv:2406.13629.
  43. 43.Kevin Wu, Eric Wu, and James Zou. 2024. Clashteval: Quantifying the tug-of-war between an llm’s internal prior and external evidence. arXiv preprint arXiv:2404.10198.
  44. 44.Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024a. Benchmarking retrieval-augmented generation for medicine. In Findings of the Association for Computational Linguistics ACL 2024, pages 6233–6251, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  45. 45.Guangzhi Xiong, Qiao Jin, Xiao Wang, Minjia Zhang, Zhiyong Lu, and Aidong Zhang. 2024b. Improving retrieval-augmented generation in medicine with iterative follow-up questions. arXiv preprint arXiv:2408.00727.
  46. 46.Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. arXiv preprint arXiv:2401.15884.
  47. 47.Zijun Yao, Weijian Qi, Liangming Pan, Shulin Cao, Linmei Hu, Weichuan Liu, Lei Hou, and Juanzi Li. 2024. Seakr: Self-aware knowledge retrieval for adaptive retrieval augmented generation. arXiv preprint arXiv:2406.19215.
  48. 48.Cyril Zakka, Rohan Shad, Akash Chaurasia, Alex R Dalal, Jennifer L Kim, Michael Moor, Robyn Fong, Curran Phillips, Kevin Alexander, Euan Ashley, et al. 2024. Almanac—retrieval-augmented language models for clinical medicine. NEJM AI, 1(2):AIoa2300068.
  49. 49.Tianjun Zhang, Shishir G Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E Gonzalez. 2024. Raft: Adapting language model to domain specific rag. arXiv preprint arXiv:2403.10131.
  50. 50.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  51. 51.Zihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad, and Jun Wang. 2023. How do large language models capture the ever-changing world knowledge? a review of recent advances. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8289–8311.

Citation

MLA
Sohn, J., et al. “Rationale-Guided Retrieval Augmented Generation for Medical Question Answering”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 12739–53, https://doi.org/10.18653/v1/2025.naacl-long.635.
APA
Sohn, J., Park, Y., Yoon, C., Park, S., Hwang, H., Sung, M., Kim, H., & Kang, J. (2025). Rationale-Guided Retrieval Augmented Generation for Medical Question Answering. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 12739–12753. https://doi.org/10.18653/v1/2025.naacl-long.635
Chicago
Sohn, J., Y. Park, C. Yoon, et al. 2025. “Rationale-Guided Retrieval Augmented Generation for Medical Question Answering”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 12739–53. https://doi.org/10.18653/v1/2025.naacl-long.635.
Harvard
Sohn, J. et al. (2025) “Rationale-Guided Retrieval Augmented Generation for Medical Question Answering”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12739–12753. Available at: https://doi.org/10.18653/v1/2025.naacl-long.635.
Vancouver
1. Sohn J, Park Y, Yoon C, Park S, Hwang H, Sung M, Kim H, Kang J (2025) Rationale-Guided Retrieval Augmented Generation for Medical Question Answering. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 12739–12753

BibTeX

@inproceedings{sohn-etal-2025-rationale,
    title = "Rationale-Guided Retrieval Augmented Generation for Medical Question Answering",
    author = "Sohn, Jiwoong  and
      Park, Yein  and
      Yoon, Chanwoong  and
      Park, Sihyeon  and
      Hwang, Hyeon  and
      Sung, Mujeen  and
      Kim, Hyunjae  and
      Kang, Jaewoo",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.635/",
    doi = "10.18653/v1/2025.naacl-long.635",
    pages = "12739--12753",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/