Built independently by an author, for readers. Read the story and support ChapterPal

keyword

evidence retrieval

Evidence retrieval is the process in natural language processing and information retrieval of searching, identifying, and extracting relevant textual passages, sentences, or documents from a corpus to support, refute, or explain a specific claim or query. As a critical component in systems designed for automated fact-checking, open-domain question answering, and reasoning verification, it supplies the factual material required by downstream models to evaluate truthfulness, generate accurate answers, or justify step-by-step reasoning. Unlike general search methods that prioritize broad topical relevance, evidence retrieval specifically aims to surface precise, high-utility snippets that provide the necessary proof or context to substantiate or challenge a targeted statement.

4 items

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva

OrganizationsBar-Ilan UniversityGoogleTel Aviv University

Why you should read this

Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.

Prompting language models to provide step-by-step answers (e.g., “Chain-of-Thought”) is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model’s answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io.

Added

2026-10-03

What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, Peter Szolovits

OrganizationsHuazhong University of Science and TechnologyMassachusetts Institute of Technology

Why you should read this

Introduces MedQA, a multilingual question-answering dataset derived from real medical board exams in English and Chinese that exposes the severe limitations of current natural language processing models on complex clinical reasoning.

Open domain question answering (OpenQA) tasks have been recently attracting more and more attention from the natural language processing (NLP) community. In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. We implement both rule-based and popular neural methods by sequentially combining a document retriever and a machine comprehension model. Through experiments, we find that even the current best method can only achieve 36.7\%, 42.0\%, and 70.1\% of test accuracy on the English, traditional Chinese, and simplified Chinese questions, respectively. We expect MedQA to present great challenges to existing OpenQA systems and hope that it can serve as a platform to promote much stronger OpenQA models from the NLP community in the future.

Added

2026-09-16