ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Archiki PrasadSwarnadeep SahaXiang ZhouMohit Bansal

article2023EMNLP85 citations

Introduces RECEVAL, a reference-free evaluation framework that assesses language model reasoning chains by decomposing steps into fine-grained content units to measure logical correctness with natural language inference and step utility with information theory.

Listen

Large language models frequently generate multi-step reasoning chains to solve complex problems, but evaluating the quality of these intermediate steps remains a significant challenge. Traditional evaluations focus almost exclusively on whether the model reaches the correct final answer, which risks mistaking lucky guesses or flawed reasoning shortcuts for genuine problem-solving ability. While reference-based evaluations exist, obtaining high-quality human-written reasoning chains is expensive, time-consuming, and limited by the fact that valid reasoning paths are rarely unique.

The article introduces and evaluates RECEVAL (Reasoning Chain Evaluation), a reference-free framework designed to assess the quality of natural language reasoning chains as informal proofs. RECEVAL evaluates reasoning chains along two core dimensions: correctness, ensuring that each step follows logically from its premises and prior context, and informativeness, ensuring that each step provides useful progress toward the final answer.

To conduct this evaluation without human references, the framework decomposes reasoning steps into fine-grained claim triplets called Reasoning Content Units using semantic role labeling. It evaluates correctness locally within a step and globally across preceding steps using Natural Language Inference models and information-theoretic measures. Informativeness is evaluated by measuring the usable information gain each step provides toward the final answer through Pointwise V-Information. The authors evaluated the framework across three benchmark datasets—Entailment Bank, GSM-8K, and DROP—spanning deductive, mathematical, and reading comprehension tasks, benchmarking against established metrics like ROUGE, BERTScore, BARTScore, and ROSCOE.

The analysis yielded several key findings. First, RECEVAL significantly outperforms existing reference-free metrics in detecting specific reasoning errors; on Entailment Bank challenge tasks, it boosted error detection correlation from 0.62 to 0.89 for hallucinations and from 0.22 to 0.39 for swap errors. Second, it delivered higher correlations with human judgments on overall chain quality and coherence, improving quality correlation from 0.28 to 0.36 on GSM-8K. Third, high-quality human-written chains consistently exhibited positive information gain across steps, whereas uninformative chains showed noticeable drops. Finally, using RECEVAL to rerank and select model-generated candidate rationales improved downstream question-answering accuracy on GSM-8K by 3.2 percentage points over standard greedy decoding.

These findings demonstrate that assessing intermediate reasoning quality independently of final answers provides a more reliable and faithful measure of model reasoning. For organizations deploying generative language models in high-stakes environments, this approach mitigates the risk of models reaching right conclusions through wrong logic. The framework also offers an automated, reference-free verification mechanism that can filter low-quality outputs and enhance downstream application performance without requiring costly human annotations.

Organizations evaluating or deploying multi-step language models should adopt reference-free, step-level verification combining logical correctness and information gain. When reference chains are available for training, smaller fine-tuned models like T5-large are recommended for informativeness calculations; when references are unavailable, larger pre-trained models or prompted frontier models can serve as practical drop-in alternatives.

The framework operates under the assumption that the knowledge necessary to evaluate reasoning steps is largely present within the context or captured by pre-trained base models, meaning performance may degrade on tasks requiring deep, implicit domain knowledge. Additionally, the metrics do not directly target arithmetic computation errors, which are better handled via external tools like calculators. Within these stated boundary conditions, the empirical evidence provides strong confidence in RECEVAL’s ability to accurately score multi-step reasoning chains.

arXiv: 2304.10703archiki/ReCEval
Cover for ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Abstract

Multi-step reasoning ability is fundamental to many natural language tasks, yet it is unclear what constitutes a good reasoning chain and how to evaluate them. Most existing methods focus solely on whether the reasoning chain leads to the correct conclusion, but this answer-oriented view may confound reasoning quality with other spurious shortcuts to predict the answer. To bridge this gap, we evaluate reasoning chains by viewing them as informal proofs that derive the final answer. Specifically, we propose RECEVAL (Reasoning Chain Evaluation), a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, i.e., each step provides new information that is helpful towards deriving the generated answer. We evaluate these properties by developing metrics using natural language inference models and V-Information. On multiple datasets, we show that RECEVAL effectively identifies various error types and yields notable improvements compared to prior methods. We analyze the impact of step boundaries, and previous steps on evaluating correctness and demonstrate that our informativeness metric captures the expected flow of information in high-quality reasoning chains. Finally, we show that scoring reasoning chains based on RECEVAL improves downstream task performance.1

Table of Contents

  • 1. Introduction
  • 2. Reasoning Chains: Preliminaries
  • 3. Properties of Good Reasoning Chains
  • 4. RECEVAL: Evaluation Metrics
  • 4.1 Evaluation of Intra-Step Correctness
  • 4.2 Evaluation of Inter-Step Correctness
  • 4.3 Evaluation of Informativeness
  • 4.4 RECEVAL: Overall Algorithm
  • 5. Meta-Evaluation Setup
  • 5.1 Meta-Evaluation: Datasets
  • 5.2 Meta-Evaluation: Baselines
  • 5.3 Meta-Evaluation: Correlation Measure
  • 6. Results and Discussion
  • 6.1 Effectiveness of RECEVAL
  • 6.2 Analysis of RECEVAL Metrics
  • 6.3 Utilizing RECEVAL for Evaluating and Improving Downstream Tasks
  • 7. Related Work
  • 8. Conclusion
  • Acknowledgements
  • Limitations
  • References
  • A. RECEVAL: Background and Details
  • B. Datasets and Errors
  • B.1 Entailment Bank
  • B.2 GSM-8K
  • C. Additional RECEVAL Meta-Evaluation
  • D. RECEVAL Correctness Metrics
  • E. Informativeness and Approximately Positive Information Gain (API)

Knowls

  1. Knowl 1 — RECEVAL evaluates reasoning chains by correctness and informativeness

    model/method

    RECEVAL treats a reasoning chain as an informal proof for a predicted answer. Given input context XX, reasoning chain R={s(1),…,s(n)}R=\{s^{(1)},\ldots,s^{(n)}\}, and predicted answer a^\hat a, it evaluates two complementary properties:

    • Correctness: every reasoning step must make a valid inference. A step must be supported both by claims inside that step and by the input context and conclusions established in earlier steps.
    • Informativeness: every step should add information that helps derive a^\hat a; redundant, irrelevant, or merely paraphrased steps therefore receive low informativeness scores.

    Correctness is evaluated locally within a step and globally against earlier information. Informativeness is evaluated by measuring how much adding a step improves prediction of the final answer. The framework is reference-free: it does not require a human-written reasoning chain for the evaluated instance.

  2. Knowl 2 — Reasoning Content Units enable fine-grained correctness evaluation

    definition

    A Reasoning Content Unit (RCU) is a fine-grained claim extracted from a reasoning step. For step s(i)s^{(i)}, RECEVAL separates the RCUs into premise units RCUp(i)={RCUpj(i)}j=1tRCU_p^{(i)}=\{RCU_{pj}^{(i)}\}_{j=1}^{t} and one conclusion unit RCUc(i)RCU_c^{(i)}. The premise units provide the information used by the step, while the conclusion unit is the claim inferred by the step.

    RCUs allow RECEVAL to distinguish an invalid conclusion from valid information appearing elsewhere in the same sentence. They are extracted from semantic-role-labeling subject–verb–object frames, with overlapping frames reduced to a non-overlapping subset. Premise and conclusion roles are assigned using sentence position and subordinating conjunctions such as “so,” “because,” and “since.”

  3. Knowl 3 — NLI metrics measure intra-step and inter-step correctness

    equation

    For reasoning step s(i)s^{(i)}, the entailment-based intra-step correctness score is

    intra-correct⁡entail(i)=Pentail ⁣(RCUp(i);RCUc(i)),\operatorname{intra\text{-}correct}^{(i)}_{\mathrm{entail}}=P_{\mathrm{entail}}\!\left(RCU_p^{(i)};RCU_c^{(i)}\right),

    where all premise RCUs in RCUp(i)RCU_p^{(i)} are concatenated, RCUc(i)RCU_c^{(i)} is the conclusion RCU, and PentailP_{\mathrm{entail}} is the entailment probability produced by a natural-language-inference model. RECEVAL uses strict entailment, so a conclusion that is merely neutral or unsupported by the premises receives a low score.

    Inter-step correctness measures whether the current conclusion contradicts the input context or any earlier conclusion. Let XX be the input context and let rr range over X∪{RCUc(j):1≤j<i}X\cup\{RCU_c^{(j)}:1\leq j<i\}. The score is

    inter-correct⁡(i)=1−max⁡r Pcontr ⁣(r;RCUc(i)),\operatorname{inter\text{-}correct}^{(i)}=1-\max_{r}\,P_{\mathrm{contr}}\!\left(r;RCU_c^{(i)}\right),

    where PcontrP_{\mathrm{contr}} is the contradiction probability from the NLI model. Only earlier conclusion RCUs are used, because premise RCUs overlap substantially with the input context. Both correctness scores lie in [0,1][0,1], with 11 indicating no detected entailment or contradiction problem, respectively.

  4. Knowl 4 — PVI measures conclusion support and answer-directed information gain

    equation

    RECEVAL uses pointwise V-information (PVI) to measure usable information under a chosen predictive-model family V\mathcal V. For an input string xx, target string yy, and conditioning string zz, let gg predict yy from zz and xx, and let g′g' predict yy from zz alone. The conditional PVI is

    PVI⁡(x→y∣z)=−log⁡g[z]′(y)+log⁡g[z,x](y).\operatorname{PVI}(x\rightarrow y\mid z)=-\log g'_{[z]}(y)+\log g_{[z,x]}(y).

    A positive value means that adding xx makes yy easier for the predictor to generate; a negative value means that it makes prediction harder. The PVI-based intra-step score is

    intra-correct⁡PVI(i)=PVI⁡ ⁣(RCUp(i)→RCUc(i)),\operatorname{intra\text{-}correct}^{(i)}_{\mathrm{PVI}}=\operatorname{PVI}\!\left(RCU_p^{(i)}\rightarrow RCU_c^{(i)}\right),

    which relaxes strict entailment by measuring how easily the conclusion can be generated from the premise information. The answer-directed information-gain score is

    info-gain⁡PVI(i)=PVI⁡ ⁣(s(i)→a^∣s(<i)),\operatorname{info\text{-}gain}^{(i)}_{\mathrm{PVI}}=\operatorname{PVI}\!\left(s^{(i)}\rightarrow\hat a\mid s^{(<i)}\right),

    where s(<i)s^{(<i)} denotes all preceding reasoning steps and a^\hat a is the predicted answer. A large positive value indicates that step s(i)s^{(i)} makes the answer easier to predict, whereas a near-zero or negative value indicates redundancy or harmful information. The main implementation fine-tunes T5-large on gold reasoning chains to estimate these probabilities.

  5. Knowl 5 — RECEVAL extracts RCUs and aggregates step scores conservatively

    algorithm

    RECEVAL takes an input context XX, reasoning chain R={s(1),…,s(n)}R=\{s^{(1)},\ldots,s^{(n)}\}, and predicted answer a^\hat a, and returns chain-level intra-correctness, inter-correctness, and informativeness scores.

    RCU extraction uses a 345M-parameter BERT-based semantic-role-labeling model. The system sorts the overlapping semantic frames extracted from each sentence by length, selects a disjoint subset, removes modifier frames that contain verbs already represented as separate frames, and assigns premise/conclusion roles using conjunction and position rules. The NLI component is the DeBERTa-v3-large MNLI/FEVER/ANLI/WANLI model. PVI uses a 770M-parameter T5-large model fine-tuned with learning rate 3×10−53\times10^{-5} for 10 epochs and weight decay 0.10.1; the checkpoint with the lowest development-set ROUGE-L is selected.

    Input: Context X, reasoning chain R, predicted answer a_hat
    Output: Chain scores score_intra, score_inter, score_info
    for each step s^(i) in R do
        Extract premise RCUs RCU_p^(i) and conclusion RCU_c^(i)
        score_intra^(i) <- intra-step correctness of RCU_p^(i) and RCU_c^(i)
        score_inter^(i) <- contradiction-based consistency of RCU_c^(i) with X and earlier conclusions
        score_info^(i) <- PVI information gain of s^(i) toward a_hat conditioned on earlier steps
    end for
    score_intra <- minimum of score_intra^(i) over all steps
    score_inter <- minimum of score_inter^(i) over all steps
    score_info <- minimum of score_info^(i) over all steps
    return score_intra, score_inter, score_info

    The minimum aggregation reflects the assumption that a reasoning chain is only as good as its least correct or least informative step.

  6. Knowl 6 — Approximately Positive Information Gain captures multi-step information flow

    theoretical result

    RECEVAL introduces APIkAPI_k to account for the fact that a high-quality reasoning chain need not make every individual step more informative than the preceding step. For a reasoning chain RR and positive integer window length kk, define

    APIk(R)={1,if ∑j=ii+k−1info-gain⁡PVI(j)>0 for every valid window i,0,otherwise.API_k(R)= \begin{cases} 1,&\text{if }\displaystyle\sum_{j=i}^{i+k-1}\operatorname{info\text{-}gain}^{(j)}_{\mathrm{PVI}}>0\text{ for every valid window }i,\\ 0,&\text{otherwise.} \end{cases}

    Thus, APIk(R)=1API_k(R)=1 means that every contiguous block of kk steps contributes positive total information toward the predicted answer. On the Entailment Bank challenge development set, the percentages of chains satisfying this condition were:

    Chain type API1_1 API2_2 API3_3
    Uninformative (REP) 36.4 69.4 80.7
    Uninformative (PAR) 35.3 70.5 81.4
    Uninformative (RED) 38.6 73.4 82.8
    Gold 72.7 87.7 92.0

    The results show that gold reasoning chains much more often exhibit positive information flow, especially when information is assessed over two or three consecutive steps.

  7. Knowl 7 — Meta-evaluation tests error detection without gold reasoning references

    experimental setup

    RECEVAL is evaluated as a reference-free metric by correlating its chain scores with known or human-annotated reasoning errors. The evaluation uses three datasets:

    • Entailment Bank (EB): deductive multi-step chains with programmatically introduced hallucination (HALL), negation (NEG), swap (SWAP), repetition (REP), paraphrase (PAR), and irrelevant-information (RED) errors. The authors create an EB-challenge split by perturbing intermediate inferences and avoiding lexical-overlap shortcuts.
    • GSM-8K: grade-school mathematical problems with model-generated chain-of-thought rationales and human annotations for quality, coherence, commonsense, factuality, hallucination, redundancy, repetition, logic, and arithmetic errors.
    • DROP: discrete reasoning over passages, with human annotations for quality, coherence, commonsense, factuality, hallucination, redundancy, repetition, and logic errors.

    Baselines include ROUGE-2, BERTScore, BARTScore, CTC, and ROSCOE semantic-alignment, semantic-similarity, logical-inference, and repetition metrics. For a metric score SS and binary error-status variable E∈{0,1}E\in\{0,1\}, the paper reports Somers' DD:

    DS,E=τ(E,S)τ(E,E),D_{S,E}=\frac{\tau(E,S)}{\tau(E,E)},

    where τ\tau is Kendall's tau coefficient. Higher correlation is interpreted as better ability to identify the relevant error or quality judgment.

  8. Knowl 8 — RECEVAL outperforms reference-free baselines across error types

    data/table

    The main meta-evaluation results report Somers' DD on the EB-challenge test set, GSM-8K test set, and DROP test set. Higher values indicate stronger association with the annotated error or quality measure. RECEVAL is especially strong for hallucination, negation, paraphrase, redundancy, factuality, and overall quality.

    EB correctness EB informativeness
    Metric HALL NEG SWAP REP PAR RED
    ROSCOE-SA 0.62 0.40 0.22 0.83 0.64 0.51
    ROSCOE-SS 0.34 0.40 0.09 0.81 0.62 0.54
    ROSCOE-LI 0.20 0.82 0.16 – – –
    RECEVAL 0.89 0.88 0.39 0.66 0.68 0.67
    QUAL COH COM FACT HALL RED REP LOGIC MATH
    ROSCOE-LI 0.28 0.26 0.18 0.34 0.22 0.35 0.98 0.22 0.09
    RECEVAL-correctness 0.36 0.31 0.21 0.37 0.28 0.40 0.63 0.25 0.24
    RECEVAL-informativeness 0.30 0.29 0.19 0.26 0.26 0.55 0.87 0.21 0.32
    QUAL COH COM FACT HALL RED REP LOGIC
    ROSCOE-SS 0.11 0.36 0.46 0.22 0.16 0.80 0.91 0.05
    ROSCOE-LI 0.20 0.24 0.46 0.39 -0.01 0.08 0.70 0.01
    RECEVAL-correctness 0.22 0.32 0.52 0.54 0.49 0.21 -0.12 0.16
    RECEVAL-informativeness 0.20 0.36 0.14 0.51 0.48 0.83 0.89 0.12

    On EB-challenge, RECEVAL raises hallucination correlation from 0.62 with ROSCOE-SA to 0.89 and paraphrase and redundancy correlations to 0.68 and 0.67. On GSM-8K it improves overall quality over ROSCOE-LI from 0.28 to 0.36, and on DROP it improves redundancy detection over ROSCOE-SS from 0.80 to 0.83 while matching or exceeding baseline performance on most other measures.

  9. Knowl 9 — RCU decomposition, complete history, and sentence-level steps are important

    empirical result

    Ablations on the EB-challenge development set show that RECEVAL depends on appropriate claim decomposition and step boundaries. The reported values are Somers' DD for hallucination (HALL), negation (NEG), and swap (SWAP) detection.

    Intra-correct Inter-correct
    RCU setting HALL NEG SWAP HALL NEG SWAP
    Without RCUs – – – 0.12 0.83 0.11
    Automatically identified RCUs 0.71 0.84 0.37 0.14 0.90 0.16
    Gold RCUs 0.89 0.94 0.54 0.16 0.96 0.16

    Using all preceding steps for inter-step correctness is better than using only the nearest one or two steps:

    History used HALL NEG SWAP
    k=1k=1 preceding step 0.08 0.79 0.14
    k=2k=2 preceding steps 0.10 0.84 0.17
    All preceding steps 0.14 0.90 0.16

    Sentence-level boundaries are also substantially better than treating each RCU or the entire chain as one step:

    Step boundary HALL NEG SWAP
    Each RCU 0.46 0.87 0.28
    Each sentence 0.86 0.90 0.38
    Entire chain 0.17 0.32 0.13

    These results indicate that fine-grained RCUs are needed to identify local inference failures, while sentence-level steps preserve enough context for reliable multi-step evaluation.

  10. Knowl 10 — RECEVAL scores improve downstream chain selection

    empirical result

    The authors test whether RECEVAL scores are useful beyond evaluation by selecting among sampled reasoning chains. For each GSM-8K problem, FLAN T5-XXL generates 20 chains with temperature 0.70.7. Chains are ranked using RECEVAL or ROSCOE metrics, and the selected chain supplies the final answer.

    Selection method Accuracy (%)
    Greedy decoding 17.3
    Sampling + ROSCOE (LI) 19.0
    Sampling + ROSCOE (SA, SS) 17.8
    Sampling + ROSCOE (REP) 18.6
    Sampling + RECEVAL (correctness) 19.6
    Sampling + RECEVAL (informativeness) 18.7
    Sampling + RECEVAL (both) 20.5

    Using both RECEVAL correctness and informativeness raises accuracy from 17.3% with greedy decoding to 20.5%, a 3.2-percentage-point improvement. Correctness alone gives 19.6% and informativeness alone gives 18.7%; the best ROSCOE configuration reaches 19.0%.

Coverage note — The paper's discussion of implicit-knowledge dependence and the possibility that generated chains are not faithful records of internal model reasoning was omitted as a limitation rather than a core method or result.

References

  1. 1.Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. Explanations for CommonsenseQA: New Dataset and Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3050–3065, Online. Association for Computational Linguistics.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, and Swabha Swayamdipta. 2022. Rev: Information-theoretic evaluation of free-text rationales. arXiv preprint arXiv:2210.04982.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Miruna-Adriana Clinciu, Arash Eshghi, and Helen Hastie. 2021. A study of automatic metrics for the evaluation of natural language explanations. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2376–2387, Online. Association for Computational Linguistics.
  9. 9.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  10. 10.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. arXiv preprint arXiv:2208.14271.
  11. 11.Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021. Explaining answers with entailment trees. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7358–7370, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota.
  14. 14.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  15. 15.Nan Duan, Duyu Tang, and Ming Zhou. 2020. Machine reasoning: Technology, dilemma and future. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts, pages 1–6, Online. Association for Computational Linguistics.
  16. 16.Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022. Understanding Dataset Difficulty with V-Usable Information. In International Conference on Machine Learning, pages 5988–6008. PMLR.
  17. 17.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023a. GPTScore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  18. 18.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2023b. Complexity-based prompting for multi-step reasoning. In The Eleventh International Conference on Learning Representations.
  19. 19.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 1–6, Melbourne, Australia. Association for Computational Linguistics.
  20. 20.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  21. 21.Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. ROSCOE: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations.
  22. 22.Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, et al. 2022. FOLIO: Natural language reasoning with first-order logic. arXiv preprint arXiv:2209.00840.
  23. 23.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. Advances in Neural Information Processing Systems.
  24. 24.John Hewitt, Kawin Ethayarajh, Percy Liang, and Christopher Manning. 2021. Conditional probing: measuring usable information beyond a baseline. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1626–1639, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403.
  26. 26.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Cosmos QA: Machine reading comprehension with contextual commonsense reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2391–2401, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Albert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Timothee Lacroix, Jiacheng Liu, Wenda Li, Mateja Jamnik, Guillaume Lample, and Yuhuai Wu. 2023. Draft, sketch, and prove: Guiding formal theorem provers with informal proofs. In The Eleventh International Conference on Learning Representations.
  28. 28.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Adam Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  29. 29.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  30. 30.Moritz Laurer, W v Atteveldt, Andreu Casas, and Kasper Welbers. 2022. Less annotating, more classifying–addressing the data scarcity issue of supervised machine learning with deep transfer learning and bert-nli.
  31. 31.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  32. 32.Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019. Reasoning over paragraph effects in situations. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 58–62, Hong Kong, China. Association for Computational Linguistics.
  33. 33.Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su. 2023. Deductive verification of chain-of-thought reasoning. arXiv preprint arXiv:2306.03872.
  34. 34.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  35. 35.Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023. Faithful chain-of-thought reasoning. arXiv preprint arXiv:2301.13379.
  36. 36.Ani Nenkova and Rebecca Passonneau. 2004. Evaluating content selection in summarization: The pyramid method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 145–152, Boston, Massachusetts, USA. Association for Computational Linguistics.
  37. 37.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  38. 38.Vishakh Padmakumar and He He. 2021. Unsupervised extractive summarization using pointwise mutual information. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2505–2512, Online. Association for Computational Linguistics.
  39. 39.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  40. 40.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  41. 41.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  42. 42.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  43. 43.Abulhair Saparov and He He. 2023. Language models can (kind of) reason: A systematic formal analysis of chain-of-thought. In International Conference on Learning Representations.
  44. 44.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  45. 45.Claude E Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
  46. 46.Ori Shapira, David Gabay, Yang Gao, Hadar Ronen, Ramakanth Pasunuru, Mohit Bansal, Yael Amsterdamer, and Ido Dagan. 2019. Crowdsourcing lightweight pyramids for manual summary evaluation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 682–687, Minneapolis, Minnesota. Association for Computational Linguistics.
  47. 47.Peng Shi and Jimmy Lin. 2019. Simple bert models for relation extraction and semantic role labeling. arXiv preprint arXiv:1904.05255.
  48. 48.Robert H Somers. 1962. A new asymmetric measure of association for ordinal variables. American sociological review, pages 799–811.
  49. 49.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  50. 50.Brian Thompson and Matt Post. 2020. Automatic machine translation evaluation in many languages via zero-shot paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 90–121, Online. Association for Computational Linguistics.
  51. 51.Jidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao, Hao He, and Yaohui Jin. 2021. Diagnosing the first-order logical reasoning ability through LogicNLI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3738–3747, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  52. 52.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  53. 53.Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv preprint arXiv:2305.04388.
  54. 54.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  55. 55.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  56. 56.Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho. 2021. Naturalproofs: Mathematical theorem proving in natural language. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  57. 57.Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, and Yejin Choi. 2022. Naturalprover: Grounded mathematical proof generation with language models. In Advances in Neural Information Processing Systems.
  58. 58.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  59. 59.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  60. 60.Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. 2020. A theory of usable information under computational constraints. International Conference on Learning Representations.
  61. 61.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34:27263–27277.
  62. 62.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  63. 63.Shiyue Zhang and Mohit Bansal. 2021. Finding a balanced degree of automation for summary evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6617–6632, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  64. 64.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with bert. In International Conference on Learning Representations.
  65. 65.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China. Association for Computational Linguistics.

Citation

MLA
Prasad, A., et al. “ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10066–86, https://doi.org/10.18653/v1/2023.emnlp-main.622.
APA
Prasad, A., Saha, S., (周翔), X. Z., & Bansal, M. (2023). ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10066–10086. https://doi.org/10.18653/v1/2023.emnlp-main.622
Chicago
Prasad, A., S. Saha, X. Z. (周翔), and M. Bansal. 2023. “ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10066–86. https://doi.org/10.18653/v1/2023.emnlp-main.622.
Harvard
Prasad, A. et al. (2023) “ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10066–10086. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.622.
Vancouver
1. Prasad A, Saha S, (周翔) XZ, Bansal M (2023) ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10066–10086

BibTeX

@inproceedings{prasad-etal-2023-receval,
    title = "{R}e{CE}val: Evaluating Reasoning Chains via Correctness and Informativeness",
    author = "Prasad, Archiki  and
      Saha, Swarnadeep  and
      Zhou, Xiang  and
      Bansal, Mohit",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.622/",
    doi = "10.18653/v1/2023.emnlp-main.622",
    pages = "10066--10086"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/