A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Alon JacoviYonatan BittonBernd BohnetJonathan HerzigOr HonovichMichael TsengMichael CollinsRoee AharoniMor Geva
Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.
Modern artificial intelligence systems increasingly rely on step-by-step reasoning, known as chain-of-thought prompting, to solve complex questions. While generating intermediate steps improves final answers, these reasoning chains frequently contain factual fabrications or logical fallacies that undermine their reliability. Automated verification tools are being developed to detect these errors, yet progress has been severely constrained by the lack of rigorous, step-level benchmarks that evaluate whether verifiers themselves are accurate.
The article addresses this gap by creating and evaluating REVEAL (Reasoning Verification Evaluation), a benchmark specifically designed to assess automatic verifiers of complex reasoning chains. The primary objective is to establish a rigorous evaluation framework that tests how well automated models can verify the relevance, factual attribution against external sources, and logical correctness of individual reasoning steps.
To construct this benchmark, the researchers collected 1,002 reasoning chains containing 3,360 individual steps across four diverse open-domain question-answering datasets, utilizing three major language models. The steps were annotated by a human pool across two decoupled tasks: verifying logical inference independently of factual truth, and verifying factual claims against up to three retrieved Wikipedia passages. The dataset was divided into a core evaluation benchmark of high-agreement cases and an open collection of ambiguous, borderline cases. The researchers then benchmarked several leading automated verifiers, including natural language inference classifiers and large language models.
The evaluation revealed several critical findings. First, reasoning chains generated by current language models are highly prone to error: only 20% of full chains were completely correct, with 77.3% exhibiting factual attribution failures and 18.5% containing logical errors. Second, existing automated verifiers struggle significantly, particularly with logic. While verifiers performed moderately well at detecting factual support against clear evidence, they exhibited severe biases on logical validation, frequently failing to identify flawed logic with performance dropping as low as 32% to 47% F1 on incorrect steps. Third, breaking the verification process into a pipeline that inspects individual steps before judging the whole chain substantially outperformed single-prompt whole-chain verification, improving overall correctness macro-F1 from 36–62% up to 54–76% across models.
These findings have immediate implications for system safety, risk management, and the deployment of AI in high-stakes environments. They demonstrate that organizations cannot rely on current automated verifiers as turnkey safety filters, especially when validating complex multi-step logic. The persistent gap between generation capability and verification capability creates substantial operational risk if AI systems are deployed autonomously without human oversight or specialized architectures.
Organizations developing or deploying complex reasoning systems should transition from monolithic verification prompts to modular, step-level verification pipelines. Future work must focus on improving logical inference verification, refining retrieval systems to better capture implicit common sense, and expanding benchmarks beyond Wikipedia-centric formats. Given that the core benchmark relies on high-agreement human annotations and fixed evidence passages, confidence in the findings is high for standard factual and logical tasks, though stakeholders should exercise caution when extrapolating these verifier baselines to highly specialized or estimation-heavy domains.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Read this foundational account of chain-of-thought prompting first to understand the reasoning chains that REVEAL evaluates step by step.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). Its reference-free evaluation of reasoning-chain correctness and informativeness introduces a closely related framework that helps contextualize REVEAL’s step-level benchmark.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Its formal analysis of where chain-of-thought deductions fail provides useful grounding for REVEAL’s separate evaluation of logical errors.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). Building on step-level verification, this work models how each reasoning step affects the likelihood of eventual success rather than scoring steps in isolation.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). It extends the evaluation of automated verifiers to judges comparing subtly flawed and correct responses across reasoning, mathematics, and coding.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). It continues the study of reasoning verification by training generative verifiers to produce rationales and select stronger solutions.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). It carries step-level evaluation into mathematical reasoning, assessing automated detection of invalid and redundant steps.
