Built independently by an author, for readers. Read the story and support ChapterPal

keyword

medical QA datasets

Medical QA datasets are curated collections of health- and biomedicine-related questions paired with verified answers, rationales, or reference documents used to train and evaluate computational question-answering systems. These datasets encompass diverse query types and formats, including multiple-choice licensing examinations, consumer health inquiries, and open-ended clinical case questions derived from biomedical literature, clinical records, and authoritative medical resources. In artificial intelligence and natural language processing, they serve as standard benchmarks for measuring a model's ability to interpret complex medical terminology, retrieve relevant scientific evidence, perform clinical reasoning, and generate accurate, factually grounded responses while minimizing hallucinations.

1 item

Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang

OrganizationsAIGEN SciencesKorea UniversityKyung Hee University

Why you should read this

Proposes a biomedical question-answering framework that boosts accuracy by using model-generated rationales for query formulation, balancing retrieval across diverse medical corpora, and filtering out distracting context with a perplexity-trained lightweight model.

Large language models (LLM) hold significant potential for applications in biomedicine, but they struggle with hallucinations and outdated knowledge. While retrieval-augmented generation (RAG) is generally employed to address these issues, it also has its own set of challenges: (1) LLMs are vulnerable to irrelevant or unhelpful context, (2) medical queries are often not well-targeted for helpful information, and (3) retrievers are prone to bias toward the specific source corpus they were trained on. In this study, we present RAG² (Rationale-Guided RAG), a new framework for enhancing the reliability of RAG in biomedical contexts. RAG² incorporates three key innovations: a small filtering model trained on perplexity-based labels of rationales, which selectively augments informative snippets of documents while filtering out distractors; LLM-generated rationales as queries to improve the utility of retrieved snippets; a structure designed to retrieve snippets evenly from a comprehensive set of four biomedical corpora, effectively mitigating retriever bias. Our experiments demonstrate that RAG² improves the state-of-the-art LLMs of varying sizes, with improvements of up to 6.1%, and it outperforms the previous best medical RAG model by up to 5.6% across three medical question-answering benchmarks. Our code is available at https://github.com/dmis-lab/RAG2

Added

2026-09-26