ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness
Archiki PrasadSwarnadeep SahaXiang ZhouMohit Bansal
Introduces RECEVAL, a reference-free evaluation framework that assesses language model reasoning chains by decomposing steps into fine-grained content units to measure logical correctness with natural language inference and step utility with information theory.
Large language models frequently generate multi-step reasoning chains to solve complex problems, but evaluating the quality of these intermediate steps remains a significant challenge. Traditional evaluations focus almost exclusively on whether the model reaches the correct final answer, which risks mistaking lucky guesses or flawed reasoning shortcuts for genuine problem-solving ability. While reference-based evaluations exist, obtaining high-quality human-written reasoning chains is expensive, time-consuming, and limited by the fact that valid reasoning paths are rarely unique.
The article introduces and evaluates RECEVAL (Reasoning Chain Evaluation), a reference-free framework designed to assess the quality of natural language reasoning chains as informal proofs. RECEVAL evaluates reasoning chains along two core dimensions: correctness, ensuring that each step follows logically from its premises and prior context, and informativeness, ensuring that each step provides useful progress toward the final answer.
To conduct this evaluation without human references, the framework decomposes reasoning steps into fine-grained claim triplets called Reasoning Content Units using semantic role labeling. It evaluates correctness locally within a step and globally across preceding steps using Natural Language Inference models and information-theoretic measures. Informativeness is evaluated by measuring the usable information gain each step provides toward the final answer through Pointwise V-Information. The authors evaluated the framework across three benchmark datasets—Entailment Bank, GSM-8K, and DROP—spanning deductive, mathematical, and reading comprehension tasks, benchmarking against established metrics like ROUGE, BERTScore, BARTScore, and ROSCOE.
The analysis yielded several key findings. First, RECEVAL significantly outperforms existing reference-free metrics in detecting specific reasoning errors; on Entailment Bank challenge tasks, it boosted error detection correlation from 0.62 to 0.89 for hallucinations and from 0.22 to 0.39 for swap errors. Second, it delivered higher correlations with human judgments on overall chain quality and coherence, improving quality correlation from 0.28 to 0.36 on GSM-8K. Third, high-quality human-written chains consistently exhibited positive information gain across steps, whereas uninformative chains showed noticeable drops. Finally, using RECEVAL to rerank and select model-generated candidate rationales improved downstream question-answering accuracy on GSM-8K by 3.2 percentage points over standard greedy decoding.
These findings demonstrate that assessing intermediate reasoning quality independently of final answers provides a more reliable and faithful measure of model reasoning. For organizations deploying generative language models in high-stakes environments, this approach mitigates the risk of models reaching right conclusions through wrong logic. The framework also offers an automated, reference-free verification mechanism that can filter low-quality outputs and enhance downstream application performance without requiring costly human annotations.
Organizations evaluating or deploying multi-step language models should adopt reference-free, step-level verification combining logical correctness and information gain. When reference chains are available for training, smaller fine-tuned models like T5-large are recommended for informativeness calculations; when references are unavailable, larger pre-trained models or prompted frontier models can serve as practical drop-in alternatives.
The framework operates under the assumption that the knowledge necessary to evaluate reasoning steps is largely present within the context or captured by pre-trained base models, meaning performance may degrade on tasks requiring deep, implicit domain knowledge. Additionally, the metrics do not directly target arithmetic computation errors, which are better handled via external tools like calculators. Within these stated boundary conditions, the empirical evidence provides strong confidence in RECEVAL’s ability to accurately score multi-step reasoning chains.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces foundational chain-of-thought prompting in large language models, creating the multi-step reasoning outputs that ReCEval is designed to evaluate.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot step-by-step reasoning elicitation, contextualizing the format and intermediate structure of chains analyzed by ReCEval.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Presents self-consistency across sampled reasoning paths, providing key background for evaluating answer-oriented versus step-level reasoning quality.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Reveals unfaithfulness and spurious shortcuts in chain-of-thought rationales, directly motivating ReCEval's shift toward proof-based intermediate step verification.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Provides a formal analysis of local step validity versus global proof planning, directly underpinning ReCEval's formal proof perspective on reasoning chains.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Establishes standard methodologies for evaluating factual consistency via natural language inference, which ReCEval adapts to assess step-level logical correctness.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Introduces LLM-based reference-free evaluation metrics using explicit criteria and chain-of-thought judging, setting the stage for specialized reasoning-chain evaluators.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). Builds upon step-level evaluation by formalizing validity and redundancy to assess intermediate reasoning quality beyond final-answer accuracy.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Leverages automated step-level verification principles to train process reward models and reinforce reasoning chains without human annotations.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). Extends intermediate step scoring into sequential decision-making process reward models by ranking reasoning steps via Q-values.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). Applies step-by-step process-supervised reward modeling to achieve scalable alignment and easy-to-hard generalization on complex problem solving.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). Explores generative verifiers that use chain-of-thought rationale generation to score and verify step-by-step reasoning outputs.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). Operationalizes the concept of reasoning informativeness and redundancy to compress lengthy chain-of-thought steps while preserving validity.
