A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Alon JacoviYonatan BittonBernd BohnetJonathan HerzigOr HonovichMichael TsengMichael CollinsRoee AharoniMor Geva

article2024ACL94 citations

Presents REVEAL, a benchmark dataset equipped with step-level annotations for relevance, evidence attribution, and logical correctness to systematically evaluate how well automatic verifiers detect errors in language model reasoning chains.

Listen

Modern artificial intelligence systems increasingly rely on step-by-step reasoning, known as chain-of-thought prompting, to solve complex questions. While generating intermediate steps improves final answers, these reasoning chains frequently contain factual fabrications or logical fallacies that undermine their reliability. Automated verification tools are being developed to detect these errors, yet progress has been severely constrained by the lack of rigorous, step-level benchmarks that evaluate whether verifiers themselves are accurate.

The article addresses this gap by creating and evaluating REVEAL (Reasoning Verification Evaluation), a benchmark specifically designed to assess automatic verifiers of complex reasoning chains. The primary objective is to establish a rigorous evaluation framework that tests how well automated models can verify the relevance, factual attribution against external sources, and logical correctness of individual reasoning steps.

To construct this benchmark, the researchers collected 1,002 reasoning chains containing 3,360 individual steps across four diverse open-domain question-answering datasets, utilizing three major language models. The steps were annotated by a human pool across two decoupled tasks: verifying logical inference independently of factual truth, and verifying factual claims against up to three retrieved Wikipedia passages. The dataset was divided into a core evaluation benchmark of high-agreement cases and an open collection of ambiguous, borderline cases. The researchers then benchmarked several leading automated verifiers, including natural language inference classifiers and large language models.

The evaluation revealed several critical findings. First, reasoning chains generated by current language models are highly prone to error: only 20% of full chains were completely correct, with 77.3% exhibiting factual attribution failures and 18.5% containing logical errors. Second, existing automated verifiers struggle significantly, particularly with logic. While verifiers performed moderately well at detecting factual support against clear evidence, they exhibited severe biases on logical validation, frequently failing to identify flawed logic with performance dropping as low as 32% to 47% F1 on incorrect steps. Third, breaking the verification process into a pipeline that inspects individual steps before judging the whole chain substantially outperformed single-prompt whole-chain verification, improving overall correctness macro-F1 from 36–62% up to 54–76% across models.

These findings have immediate implications for system safety, risk management, and the deployment of AI in high-stakes environments. They demonstrate that organizations cannot rely on current automated verifiers as turnkey safety filters, especially when validating complex multi-step logic. The persistent gap between generation capability and verification capability creates substantial operational risk if AI systems are deployed autonomously without human oversight or specialized architectures.

Organizations developing or deploying complex reasoning systems should transition from monolithic verification prompts to modular, step-level verification pipelines. Future work must focus on improving logical inference verification, refining retrieval systems to better capture implicit common sense, and expanding benchmarks beyond Wikipedia-centric formats. Given that the core benchmark relies on high-agreement human annotations and fixed evidence passages, confidence in the findings is high for standard factual and logical tasks, though stakeholders should exercise caution when extrapolating these verifier baselines to highly specialized or estimation-heavy domains.

arXiv: 2402.00559
Cover for A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Abstract

Prompting language models to provide step-by-step answers (e.g., “Chain-of-Thought”) is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model’s answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io.

Table of Contents

  • 1 Introduction
  • 2 Formalism for Verification of Reasoning Chains
  • 3 Annotation Schema
  • 3.1 Task 1: Relevance, Type, and Logic
  • 3.2 Task 2: Relevance and Attribution
  • 4 Data Collection Process
  • 5 REVEAL
  • 5.1 Analysis of Unsupported Claims
  • 5.2 Analysis of REVEAL-Open
  • 6 Experiments
  • 6.1 Step-level Verification
  • 6.2 CoT-level Verification
  • 7 Related Work
  • 8 Conclusion
  • 9 Limitations
  • Acknowledgments
  • References
  • A REVEAL: Extended Details
  • A.1 Source Data Collection
  • A.2 Annotation Questionnaire
  • A.3 Additional Statistics
  • A.4 Analyses
  • B Experiments and Analyses
  • B.1 Experiment Details
  • B.2 Additional Results
  • B.3 Prompt Templates

Knowls

  1. Knowl 1 — Step-level definition of reasoning-chain correctness

    definition

    A reasoning chain for a question qq is an ordered sequence r=(s1,…,sn)r=(s_1,\ldots,s_n) of standalone sentences, where each step sis_i is generated conditional on earlier steps and the final step sns_n gives the answer. A verifier receives the question, chain, and external evidence ee and scores the chain’s correctness; retrieval of ee is treated as separate from verification. The benchmark separates two correctness criteria: attribution, whether claims that introduce world knowledge are supported by the supplied evidence, and logical correctness, whether an inference follows from the question and preceding steps. A chain is correct only when every relevant step is correct; irrelevant steps do not by themselves invalidate it. A logical step following an earlier incorrect logical step is labeled undefined, because its validity cannot be assessed independently of that earlier error.

  2. Knowl 2 — Human annotation schema separates evidence attribution from reasoning logic

    model/method

    REVEAL uses two annotation tasks to reduce cognitive interference between judging evidence and judging inference. In the chain-level task, annotators mark each step as relevant or irrelevant, classify it as an attribution step (introduces externally verifiable knowledge), a logical step (draws an inference), or both, and judge its logical inference as correct or incorrect. External facts are assumed correct in this task, so the logical judgment isolates inference quality. In the evidence task, annotators assess attribution steps against supplied passages as fully attributable, partially attributable, contradictory, or unsupported; relevance is also labeled from the perspective of whether the claim helps answer the question. For a step combining factual content and inference, the protocol assumes the factual content is fully attributable when judging logic, and the inference is correct when judging attribution. Annotators provide a free-text justification for each task.

  3. Knowl 3 — REVEAL corpus construction and evidence collection

    experimental setup

    REVEAL samples open-domain questions from four QA datasets: StrategyQA, two-hop MuSiQue, Sports Understanding, and Fermi; the four sources are represented approximately evenly. Three language models generate the chains: Flan-PaLM-540B, GPT-3 (text-davinci-003), and Flan-UL2-20B. One quarter of the sampled questions are answered by all three models, while the rest are distributed among the models. Generated answers are split into sentence-level steps. For attribution annotation, each step is decontextualized to make references explicit, then paired with up to three paragraphs retrieved from a 2021 Wikipedia snapshot: two from GTR dense retrieval and one from BM25. The annotation pool comprises 13 English-speaking annotators; five annotations are collected for each question–chain pair in each of the two tasks, and each label has an annotator-written justification. Attribution annotation stops when fully supporting or contradicting evidence is found, or after the available three passages have been judged.

  4. Knowl 4 — REVEAL dataset scale and split composition

    data/table

    REVEAL contains a higher-agreement evaluation split, REVEAL-Eval, and a smaller open split, REVEAL-Open, reserved for chains with at least one low-agreement step. The corpus provides hundreds of questions and thousands of reasoning steps for evaluation, together with evidence and step-level labels.

    Quantity REVEAL-Eval REVEAL-Open
    Questions 704 205
    CoT answers 1,002 224
    CoT steps 3,360 847
    Average steps per CoT 3.4 5.1
    Attribution steps 1,979 485
    Step–evidence pairs 3,502 745
    Average evidence length (words) 103 103
    Logic steps 1,250 306
    Fully attributable step–evidence pairs 864 –
    Logically correct steps 1,063 –
    Fully correct CoT answers 200 –
  5. Knowl 5 — REVEAL labels show frequent attribution failures and imperfect agreement

    empirical result

    In REVEAL-Eval, 98.6% of steps are labeled relevant, and 87.5% of logical steps are labeled logically correct. Attribution-step labels are 43.8% fully attributable, 38.6% unsupported, 11.5% contradictory, and 6.1% partially attributable. The corresponding step–evidence-pair labels are 24.7% fully attributable, 65.4% unsupported, 6.5% contradictory, and 3.4% partially attributable. The authors caution that an unsupported label is evidence-relative: it can reflect a false claim, but it can also result from a failure to retrieve or recognize supporting evidence. Krippendorff’s α\alpha is 0.49 for attribution steps and 0.46 for logical steps. A CoT answer is assigned to REVEAL-Open if any step has fewer than three of five annotators agreeing on a label; this criterion applies to 18% of CoT answers. In REVEAL-Eval, 77.3% of chains contain a step that is not fully attributable, and 18.5% contain a logically incorrect step. Full-chain correctness requires all steps to be fully attributable and logically correct; the dataset reports 200 such answers. Full-correctness rates vary by source dataset (StrategyQA 21.5%, MuSiQue 17.8%, Sports Understanding 32.4%, Fermi 3.6%) and, among the subset with all three model answers in REVEAL-Eval, by generator (Flan-UL2 11.5%, Flan-PaLM 25.0%, GPT-3 18.8%).

  6. Knowl 6 — Unsupported claims and low-agreement labels have identifiable causes

    empirical result

    A qualitative audit of 40 unsupported attribution steps (10 sampled from each source dataset) found that 19 claims were factually incorrect and 21 were judged factually correct. Among those 21 correct claims, 13 required additional reasoning or world knowledge despite relevant retrieved evidence, six had irrelevant retrieved evidence, one contained multiple claims of which only one was supported, and one was insufficiently decontextualized. The authors estimate that roughly half of unsupported cases stem from imperfect retrieval. In REVEAL-Open, disagreements were manually categorized; a step could receive multiple categories, so the percentages below do not sum to 100%. Attribution disagreements most often involved world knowledge/general inference (25.63%), rating-category criteria (23.62%), specialized knowledge (13.07%), insufficient hedging (13.07%), unclear references (12.56%), averages or ranges (10.05%), and temporal inconsistency (9.55%). For logical steps, prominent categories were calculation or unit issues (42.31%), unclear reference or standard (30.77%), invalid inference in a previous step (28.85%), relevance disputes (23.08%), and world knowledge/general inference (15.38%). These analyses show that unsupported evidence labels and annotator disagreement can arise from retrieval gaps, implicit knowledge, claim phrasing, and dependencies between steps—not only from factual or logical errors.

  7. Knowl 7 — Verifier benchmark compares step-level and whole-chain systems

    experimental setup

    The REVEAL-Eval experiments test step-level attribution, logic, and step-type classification, as well as binary correctness of an entire CoT. Baselines include few-shot Flan-UL2-20B, Flan-PaLM-540B, PaLM-2-L, and GPT-3 (text-davinci-003); a T5-XXL model with 11B parameters trained on a mixture of NLI datasets; and FacTool, a GPT-3-based fact-checking pipeline. Attribution is tested both as a binary entailment/not-entailment task and as a three-way fully supported/contradictory/not-enough-information task. The prompting baselines use five independently sampled 8-shot prompts per task, averaged for evaluation; demonstrations are class-balanced and sampled from sets of 13. FacTool is given REVEAL’s evidence rather than retrieving evidence itself. Results are reported as macro-F1 because label distributions are imbalanced. Whole-chain verification is tested with either a single prompted decision based on the question, chain, and evidence, or a pipeline that combines the step-level predictions to check whether an incorrect step occurs.

  8. Knowl 8 — Step-level verifiers struggle most with logic and step-type classification

    empirical result

    Macro-F1 for REVEAL-Eval step-level tasks shows that the NLI-trained T5 baseline leads on binary attribution, while PaLM-2-L leads among the prompting systems on three-way attribution, logic, and step type. All tested systems perform poorly on step type (all reported scores are below 65.0). The class balances are 76:24 for binary attribution, 70:24:6 for three-way attribution, 80:20 for logic, and 59:40:1 for type. Relevance results are omitted because all models collapsed to the majority-class baseline.

    Baseline Attribution 2-class Attribution 3-class Logic Type
    Flan-UL2-20B 65.2 50.4 59.4 27.3
    Flan-PaLM-540B 85.1 66.0 68.6 51.2
    PaLM-2-L 85.9 70.7 77.6 64.1
    GPT-3 81.4 51.3 59.4 52.3
    FacTool 71.1 – – –
    T5-XXL-TRUE 88.4 55.0 47.3 –
  9. Knowl 9 — Step-level pipelines outperform direct whole-chain judgments

    empirical result

    For binary CoT correctness on REVEAL-Eval, combining step-level decisions yields higher macro-F1 than prompting a model for one whole-chain decision for every tested language model. PaLM-2-L has the strongest pipeline score, 76.4 macro-F1; the pipeline advantage is especially associated with identifying incorrect chains. The task has an 80:20 class balance, with incorrect CoTs as the majority class.

    Macro-F1 Correct-CoT F1 Incorrect-CoT F1
    Baseline Single Pipeline Single Pipeline Single Pipeline
    Flan-UL2-20B 41.5 54.4 29.3 29.3 54.0 79.5
    Flan-PaLM-540B 39.4 58.1 40.2 31.9 38.6 84.3
    PaLM-2-L 61.9 76.4 51.4 65.2 72.4 87.5
    GPT-3 35.6 71.9 37.1 56.7 34.4 87.2

    The pipeline combines the model’s step-level relevance, type, attribution, and logic predictions to detect whether an incorrect step exists. Direct single-decision systems remain weak at identifying correct chains, and the results overall indicate substantial difficulty with whole-chain verification.

  10. Knowl 10 — Benchmark labels are evidence- and format-specific

    limitation

    REVEAL evaluates attribution to supplied evidence, not end-to-end fact checking with evidence retrieval; its labels therefore characterize the particular passages provided. Some claims labeled unsupported may have evidence that the retriever did not surface. Wikipedia is the evidence source, which is a particularly poor fit for Fermi questions designed to require difficult-to-source estimates and reasoning. The corpus also uses sentence-separated Chain-of-Thought answers: chains that combine multiple claims in one sentence or merge knowledge and inference may require claim extraction not tested by this benchmark.

Coverage note — The supplementary annotation-interface screenshots, verbatim prompt templates, and full illustrative examples are omitted because the knowls capture their contributed protocol, evaluation design, and findings without reproducing implementation appendices.

References

  1. 1.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023. Palm 2 technical report.
  2. 2.Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O’Reilly Media, Inc.".
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  5. 5.Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. 2023. Reconcile: Round-table conference improves reasoning via consensus among diverse llms.
  6. 6.I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu, et al. 2023. Factool: Factuality detection in generative ai–a tool augmented framework for multi-task and multi-domain scenarios. arXiv preprint arXiv:2307.13528.
  7. 7.Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021. Decontextualization: Making sentences stand-alone. Transactions of the Association for Computational Linguistics, 9:447–461.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168.
  11. 11.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The PASCAL recognising textual entailment challenge. In Machine Learning Challenges, Evaluating Predictive Uncertainty, Visual Object Classification and Recognizing Textual Entailment, First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers, volume 3944 of Lecture Notes in Computer Science, pages 177–190. Springer.
  12. 12.Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021. Explaining answers with entailment trees. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7358–7370, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  13. 13.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023. RARR: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16477–16508, Toronto, Canada. Association for Computational Linguistics.
  14. 14.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies. Transactions of the Association for Computational Linguistics (TACL).
  15. 15.Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. ROSCOE: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  16. 16.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  17. 17.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Xiaodong Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. ArXiv, abs/2103.03874.
  18. 18.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating factual consistency evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3905–3920, Seattle, United States. Association for Computational Linguistics.
  19. 19.Yushi Hu, Otilia Stretcu, Chun-Ta Lu, Krishnamurthy Viswanathan, Kenji Hata, Enming Luo, Ranjay Krishna, and Ariel Fuxman. 2023. Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.
  20. 20.Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg. 2023a. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.
  21. 21.Alon Jacovi, Avi Caciularu, Jonathan Herzig, Roee Aharoni, Bernd Bohnet, and Mor Geva. 2023b. A comprehensive evaluation of tool-assisted generation strategies. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 13856–13878, Singapore. Association for Computational Linguistics.
  22. 22.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022. Maieutic prompting: Logically consistent reasoning with recursive explanations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1266–1279, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  23. 23.Ashwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal, and Peter Clark. 2021. How much coffee was consumed during EMNLP 2019? fermi problems: A new reasoning challenge for AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7318–7328, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. Scitail: A textual entailment dataset from science question answering. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1).
  25. 25.Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2018. Towards automated factchecking: Developing an annotation schema and benchmark for consistent automated claim detection. CoRR, abs/1809.08193.
  26. 26.Andrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan, Allison C. Tam, James L. McClelland, Chen Yan, Adam Santoro, Neil C. Rabinowitz, Jane X. Wang, and Felix Hill. 2022. Tell me why! explanations support learning relational and causal structure. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 11868–11890. PMLR.
  27. 27.Christoph Leiter, Piyawat Lertvittayakumjorn, M. Fomicheva, Wei Zhao, Yang Gao, and Steffen Eger. 2022. Towards explainable evaluation metrics for natural language generation. ArXiv, abs/2203.11131.
  28. 28.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2023. Making language models better reasoners with step-aware verifier. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5315–5333, Toronto, Canada. Association for Computational Linguistics.
  29. 29.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step.
  30. 30.Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman. 2021. What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1221–1235, Online. Association for Computational Linguistics.
  31. 31.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2021. Large dual encoders are generalizable retrievers.
  32. 32.Juri Opitz and Anette Frank. 2021. Towards a decomposable metric for explainable evaluation of text generation from AMR. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1504–1518, Online. Association for Computational Linguistics.
  33. 33.Simon Ott, Konstantin Hebenstreit, Valentin Liévin, Christoffer Egeberg Hother, Milad Moradi, Maximilian Mayrhauser, Robert Praas, Ole Winther, and Matthias Samwald. 2023. Thoughtsource: A central hub for large language model reasoning data.
  34. 34.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. CoRR, abs/2105.11447.
  35. 35.Archiki Prasad, Swarnadeep Saha, Xiang Zhou, and Mohit Bansal. 2023. ReCEval: Evaluating reasoning chains via correctness and informativeness. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10066–10086, Singapore. Association for Computational Linguistics.
  36. 36.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models.
  37. 37.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. CoRR, abs/2112.12870.
  38. 38.Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. 2023. Why don’t you do it right? analysing annotators’ disagreement in subjective tasks. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2428–2441, Dubrovnik, Croatia. Association for Computational Linguistics.
  39. 39.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin C! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 624–643, Online. Association for Computational Linguistics.
  40. 40.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakaş, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartłomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, César Ferri Ramírez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodola, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovich-López, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Schütze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Kocon, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jörg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Colón, Luke Metz, Lütfi Kerem Şenel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramírez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Michał Swędrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan A. Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Miłkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan LeBras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Tushniwal, Shyam Upadhyay, Shyamolima Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsu Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.
  41. 41.Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 641–651, New Orleans, Louisiana. Association for Computational Linguistics.
  42. 42.Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. Ul2: Unifying language learning paradigms.
  43. 43.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  44. 44.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multi-hop questions via single-hop question composition. Transactions of the Association for Computational Linguistics.
  45. 45.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  46. 46.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models.
  47. 47.Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018. Constructing datasets for multi-hop reading comprehension across documents. Transactions of the Association for Computational Linguistics, 6:287–302.
  48. 48.Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, and Yejin Choi. 2022. Naturalprover: Grounded mathematical proof generation with language models. ArXiv, abs/2205.12910.
  49. 49.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  50. 50.Hugo Zaragoza, Nick Craswell, Michael Taylor, Suchi Saria, and Stephen Robertson. 2004. Microsoft cambridge at trec-13: Web and hard tracks. In IN PROCEEDINGS OF TREC 2004.
  51. 51.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022. Star: Bootstrapping reasoning with reasoning.
  52. 52.Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2023. How language model hallucinations can snowball.

Citation

MLA
Jacovi, A., et al. “A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 4615–34, https://doi.org/10.18653/v1/2024.acl-long.254.
APA
Jacovi, A., Bitton, Y., Bohnet, B., Herzig, J., Honovich, O., Tseng, M., Collins, M., Aharoni, R., & Geva, M. (2024). A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4615–4634. https://doi.org/10.18653/v1/2024.acl-long.254
Chicago
Jacovi, A., Y. Bitton, B. Bohnet, et al. 2024. “A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4615–34. https://doi.org/10.18653/v1/2024.acl-long.254.
Harvard
Jacovi, A. et al. (2024) “A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4615–4634. Available at: https://doi.org/10.18653/v1/2024.acl-long.254.
Vancouver
1. Jacovi A, Bitton Y, Bohnet B, Herzig J, Honovich O, Tseng M, Collins M, Aharoni R, Geva M (2024) A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4615–4634

BibTeX

@inproceedings{jacovi-etal-2024-chain,
    title = "A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains",
    author = "Jacovi, Alon  and
      Bitton, Yonatan  and
      Bohnet, Bernd  and
      Herzig, Jonathan  and
      Honovich, Or  and
      Tseng, Michael  and
      Collins, Michael  and
      Aharoni, Roee  and
      Geva, Mor",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.254/",
    doi = "10.18653/v1/2024.acl-long.254",
    pages = "4615--4634"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/