keyword
evidence retrieval
Evidence retrieval is the process in natural language processing and information retrieval of searching, identifying, and extracting relevant textual passages, sentences, or documents from a corpus to support, refute, or explain a specific claim or query. As a critical component in systems designed for automated fact-checking, open-domain question answering, and reasoning verification, it supplies the factual material required by downstream models to evaluate truthfulness, generate accurate answers, or justify step-by-step reasoning. Unlike general search methods that prioritize broad topical relevance, evidence retrieval specifically aims to surface precise, high-utility snippets that provide the necessary proof or context to substantiate or challenge a targeted statement.
4 items

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva
Why you should read this
Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.
Prompting language models to provide step-by-step answers (e.g., “Chain-of-Thought”) is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model’s answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io.
Added
2026-10-03

Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation
Max Glockner, Yufang Hou, Iryna Gurevych
Why you should read this
Reveals that automated fact-checking models fail on real-world misinformation because they rely on leaked counter-evidence from post-hoc reports rather than disproving the underlying reasoning behind novel claims.
The task of misinformation detection has tremendous potential to make a significant contribution to society. However, current automatic fact-checking systems ignore an important aspect of fact-checking: the lack of counter-evidence. In this paper, we analyze the prevalence of counter-evidence in real-world fact-checking datasets. We find that counter-evidence is almost always present in the evidence used by human fact-checkers, but is missing in the evidence retrieved by current automatic fact-checking systems. We argue that this discrepancy is a major reason for the lack of robustness of current systems. To address this issue, we propose a new task setting that requires the system to identify the lack of counter-evidence and to abstain from making a prediction in such cases. We show that this is a challenging task for current systems, and that our proposed method can improve the robustness of fact-checking systems.
Added
2026-10-03

What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, Peter Szolovits
Why you should read this
Introduces MedQA, a multilingual question-answering dataset derived from real medical board exams in English and Chinese that exposes the severe limitations of current natural language processing models on complex clinical reasoning.
Open domain question answering (OpenQA) tasks have been recently attracting more and more attention from the natural language processing (NLP) community. In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. We implement both rule-based and popular neural methods by sequentially combining a document retriever and a machine comprehension model. Through experiments, we find that even the current best method can only achieve 36.7\%, 42.0\%, and 70.1\% of test accuracy on the English, traditional Chinese, and simplified Chinese questions, respectively. We expect MedQA to present great challenges to existing OpenQA systems and hope that it can serve as a platform to promote much stronger OpenQA models from the NLP community in the future.
Added
2026-09-16

FEVER: a Large-scale Dataset for Fact Extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Arpit Mittal
Why you should read this
Introduces FEVER, a benchmark dataset of over 185,000 claims that challenges automated fact-checking systems to jointly verify assertions and retrieve supporting textual evidence from Wikipedia.
In this paper we introduce a new publicly available dataset for verification against textual sources, FEVER: Fact Extraction and VERification. It consists of 185,445 claims generated by altering sentences extracted from Wikipedia and subsequently verified without knowledge of the sentence they were derived from. The claims are classified as Supported, Refuted or NotEnoughInfo by annotators achieving 0.6841 in Fleiss . For the first two classes, the annotators also recorded the sentence(s) forming the necessary evidence for their judgment. To characterize the challenge of the dataset presented, we develop a pipeline approach and compare it to suitably designed oracles. The best accuracy we achieve on labeling a claim accompanied by the correct evidence is 31.87%, while if we ignore the evidence we achieve 50.91%. Thus we believe that FEVER is a challenging testbed that will help stimulate progress on claim verification against textual sources.
Added
2026-09-14
