Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiple-choice QA

Multiple-choice question answering is a natural language processing task in which an automated system is presented with a question and a predefined set of candidate options, with the goal of selecting the correct answer from among plausible distractors. Unlike extractive or free-form generative question answering, which requires locating text spans or generating unstructured responses, multiple-choice question answering evaluates a model by framing the problem as ranking or classification across a closed set of alternatives. This format serves as a standard benchmark for measuring artificial intelligence capabilities across diverse areas, including reading comprehension, factual recall, commonsense reasoning, and domain-specific expertise in subjects such as science and medicine. By limiting the valid responses to explicit options, the task facilitates objective accuracy measurements while testing a system ability to reason through complex information and filter out misleading alternatives.

2 items

Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Rationale-Guided Retrieval Augmented Generation for Medical Question Answering

Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang

OrganizationsAIGEN SciencesKorea UniversityKyung Hee University

Why you should read this

Proposes a biomedical question-answering framework that boosts accuracy by using model-generated rationales for query formulation, balancing retrieval across diverse medical corpora, and filtering out distracting context with a perplexity-trained lightweight model.

Large language models (LLM) hold significant potential for applications in biomedicine, but they struggle with hallucinations and outdated knowledge. While retrieval-augmented generation (RAG) is generally employed to address these issues, it also has its own set of challenges: (1) LLMs are vulnerable to irrelevant or unhelpful context, (2) medical queries are often not well-targeted for helpful information, and (3) retrievers are prone to bias toward the specific source corpus they were trained on. In this study, we present RAG² (Rationale-Guided RAG), a new framework for enhancing the reliability of RAG in biomedical contexts. RAG² incorporates three key innovations: a small filtering model trained on perplexity-based labels of rationales, which selectively augments informative snippets of documents while filtering out distractors; LLM-generated rationales as queries to improve the utility of retrieved snippets; a structure designed to retrieve snippets evenly from a comprehensive set of four biomedical corpora, effectively mitigating retriever bias. Our experiments demonstrate that RAG² improves the state-of-the-art LLMs of varying sizes, with improvements of up to 6.1%, and it outperforms the previous best medical RAG model by up to 5.6% across three medical question-answering benchmarks. Our code is available at https://github.com/dmis-lab/RAG2

Added

2026-09-26

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord

OrganizationsAllen Institute for AI

Why you should read this

Introduces the AI2 Reasoning Challenge (ARC), a benchmark of 7,787 natural science questions specifically filtered to defeat standard retrieval and co-occurrence methods, revealing that leading question-answering models struggle to outperform random guessing on advanced reasoning tasks.

We present a new question set, text corpus, and baselines assembled to encourage AI research in advanced question answering. Together, these constitute the AI2 Reasoning Challenge (ARC), which requires far more powerful knowledge and reasoning than previous challenges such as SQuAD or SNLI. The ARC question set is partitioned into a Challenge Set and an Easy Set, where the Challenge Set contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurence algorithm. The dataset contains only natural, grade-school science questions (authored for human tests), and is the largest public-domain set of this kind (7,787 questions). We test several baselines on the Challenge Set, including leading neural models from the SQuAD and SNLI tasks, and find that none are able to significantly outperform a random baseline, reflecting the difficult nature of this task. We are also releasing the ARC Corpus, a corpus of 14M science sentences relevant to the task, and implementations of the three neural baseline models tested. Can your model perform better? We pose ARC as a challenge to the community.

Added

2026-09-09