Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters

Boshi WangSewon MinXiang DengJiaming ShenYou WuLuke ZettlemoyerHuan Sun

article2023ACL411 citations

Reveals that chain-of-thought prompting remains highly effective even when demonstrations contain invalid reasoning, demonstrating that query relevance and step ordering drive performance gains rather than the logical validity of exemplar rationales.

Listen

Large language models have shown remarkable success on complex tasks through Chain-of-Thought prompting, a technique that supplies step-by-step example rationales within the prompt to encourage explicit multi-step reasoning. Despite its widespread adoption, practitioners and researchers have lacked a clear understanding of why this technique works and which components of the demonstrated reasoning steps are truly essential. Understanding these mechanisms is critical for designing cost-effective prompt engineering pipelines and accurately benchmarking model capabilities.

The article evaluates what makes Chain-of-Thought prompting effective by systematically altering different components of the prompt demonstrations. Specifically, the analysis demonstrates how model performance changes when the logical validity, relevance, and ordering of demonstrated rationales are intentionally degraded.

The authors conducted controlled ablation experiments on two representative multi-step reasoning benchmarks: the GSM8K mathematical reasoning dataset (evaluated on a 800-sample test subset) and the Bamboogle multi-hop factual question answering dataset (125 test questions). They tested several high-capacity models, primarily InstructGPT (text-davinci-002 and text-davinci-003), along with PaLM and Flan-PaLM. The evaluation examined both final output correctness and intermediate intrinsic quality by tracking whether models correctly derived key intermediate entities and numbers.

The key findings reveal that logical validity in prompt demonstrations matters far less than previously assumed. First, providing demonstrations with completely invalid, flawed reasoning steps still enabled models to achieve over 80% to 90% of standard Chain-of-Thought performance, and the models continued to generate coherent, logically sound reasoning during inference. Second, prompt relevance is essential; completely removing query relevance from the demonstrated objects and templates caused performance on mathematical reasoning to collapse from an intermediate score of 48.3% down to 11.9%, falling below basic standard prompting (15.4%). Third, the structural coherence and ordering of the language templates are critical, whereas maintaining the exact sequence of intermediate numbers and entities matters substantially less. Finally, models with extensive pretraining or instruction tuning on similar tasks proved highly resilient to flawed prompt rationales.

These findings imply that large language models do not primarily learn how to reason step-by-step from few-shot in-context demonstrations. Instead, the models already possess underlying reasoning capabilities acquired during large-scale pretraining. Demonstrations mainly serve as formatting guides that activate existing capabilities and enforce structured output. This insight alters operational risk assessments: while prompt engineers do not need to spend excessive effort crafting perfectly verified logical steps, models risk over-relying on pretraining priors and ignoring critical instructions or counterfactual context.

Decision-makers and engineering teams should focus prompt design resources on ensuring strict topic relevance and coherent language framing rather than exhaustive step-by-step manual verification of prompt logic. For benchmarking and evaluation, organizations must develop novel test suites where models have minimal prior exposure, ensuring tests measure true in-context skill acquisition rather than the mere recitation of pretraining data.

These conclusions should be interpreted within the context of specific limitations. The empirical analysis focused on arithmetic and factual question answering benchmarks; highly template-based symbolic tasks were not evaluated due to rigid structures. Furthermore, the experiments relied on single-run evaluations without reporting variance across multiple random seeds, and invalid reasoning steps were written manually rather than generated via a formal algorithmic taxonomy. Nonetheless, the consistent trends observed across diverse model architectures provide high confidence that prompt relevance and structural ordering outweigh step-by-step logical validity in eliciting model reasoning.

Cover for Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters

Abstract

Chain-of-Thought (CoT) prompting can dramatically improve the multi-step reasoning abilities of large language models (LLMs). CoT explicitly encourages the LLM to generate intermediate rationales for solving a problem, by providing a series of reasoning steps in the demonstrations. Despite its success, there is still little understanding of what makes CoT prompting effective and which aspects of the demonstrated reasoning steps contribute to its performance. In this paper, we show that CoT reasoning is possible even with invalid demonstrations—prompting with invalid reasoning steps can achieve over 80-90% of the performance obtained using CoT under various metrics, while still generating coherent lines of reasoning during inference. Further experiments show that other aspects of the rationales, such as being relevant to the query and correctly ordering the reasoning steps, are much more important for effective CoT reasoning. Overall, these findings both deepen our understanding of CoT prompting, and open up new questions regarding LLMs’ capability to learn to reason in context.1

Table of Contents

  • 1 Introduction
  • 2 Background & Study Formulation
  • Arithmetic Reasoning
  • Multi-hop QA
  • 3 Experimental Setup
  • 3.1 Datasets & In-context Exemplars
  • 3.2 Backbone Language Model
  • 3.3 Evaluation
  • 4 How Much Does Valid Reasoning Matter?
  • 4.1 Constructing Invalid Chain of Reasoning
  • 4.2 Results & Analysis
  • 5 What are the Key Aspects of Chain-of-Thoughts?
  • 5.1 Ablation Settings
  • 5.2 Results & Analysis
  • 6 Discussion
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Chain of Thought Exemplars
  • A.2 More Details on Intrinsic Evaluation
  • A.3 Additional Results & Discussion
  • A.4 Full List of Prompts

Knowls

  1. Knowl 1 — Robustness of Multi-Step Reasoning to Invalid In-Context Reasoning Demonstrations

    empirical result

    When large language models (LLMs) are evaluated on multi-step reasoning tasks using Chain-of-Thought (CoT) prompting, replacing the valid reasoning steps in the few-shot demonstration exemplars with completely invalid and fallacious reasoning steps results in only a marginal decrease in performance.

    Specifically, on arithmetic reasoning (GSM8K) and multi-hop factual question answering (Bamboogle) using models such as InstructGPT (text-davinci-002 and text-davinci-003), PaLM, and Flan-PaLM, prompting with invalid reasoning achieves over 80% to 90% of the extrinsic performance (exact-match answer accuracy / F1) and over 90% of the intrinsic intermediate reasoning recall/F1 obtained by standard valid CoT prompting. Furthermore, across different problem difficulty levels (measured by reasoning depth in GSM8K), the performance drop remains uniformly small.

    This demonstrates that the logical validity of the intermediate steps in few-shot demonstrations contributes only a minor portion to the overall effectiveness of CoT prompting, indicating that LLMs do not learn how to reason step-by-step from in-context exemplars, but rather utilize demonstrations to activate existing reasoning capabilities acquired during pretraining.

  2. Knowl 2 — Structural Decomposition of Chain-of-Thought Rationales: Bridging Objects and Language Templates

    definition

    A Chain-of-Thought (CoT) rationale is partitioned into two complementary structural components:

    1. Bridging Objects: The key, essential intermediate values or entities that a model must traverse to reach the final answer. In arithmetic reasoning, bridging objects are defined as the numerical values and intermediate equations (e.g., intermediate quantities computed across calculation steps). In multi-hop factual question answering, bridging objects are intermediate subject and object entities connecting the question to the target entity.

    2. Language Templates: The complementary textual scaffolding surrounding bridging objects. These consist of natural language descriptions, relations, and predicates that guide the derivation and contextual interpretation of bridging objects along the multi-step reasoning path.

  3. Knowl 3 — Dimensions of Chain-of-Thought Rationale Integrity: Relevance and Coherence

    definition

    The structural validity of Chain-of-Thought rationales is characterized along two orthogonal dimensions:

    1. Relevance: A rationale component is relevant if it is directly grounded in and pertains to the input query. For bridging objects, relevance requires utilizing the exact numerical values or named entities specified in the query. For language templates, relevance requires preserving the specific relations, entities, and question topic introduced by the query.

    2. Coherence: A rationale component is coherent if its constituent steps follow an ordered, logically chronological progression such that subsequent steps depend on prior steps rather than preceding their own preconditions. For bridging objects, coherence requires that operations operate on previously introduced values. For language templates, coherence requires that narrative dependencies and relational clauses occur in valid causal order.

  4. Knowl 4 — Empirical Comparison of Chain-of-Thought Ablation Configurations on InstructGPT

    data/table

    The following table reports the performance of InstructGPT (text-davinci-002) evaluated under greedy decoding (T=0T = 0) across standard few-shot prompting (STD), Chain-of-Thought prompting (CoT), invalid reasoning prompting, and selective ablations of relevance and coherence across bridging objects and language templates. Evaluations are performed on GSM8K (800 test samples) and Bamboogle (125 test samples).

    Prompt Setting GSM8K Bamboogle
    Inter. Recall Inter. F1 Answer Acc. Inter. Recall Answer F1
    STD (Standard prompting) N/A N/A 15.4 N/A 20.6
    CoT (Chain-of-Thought) 43.9 48.3 48.5 45.2 45.2
    (1) Invalid Reasoning 39.8 43.9 39.5 44.4 39.4
    (2) No coherence for bridging objects 35.3 39.2 35.8 40.8 37.4
    (3) No relevance for bridging objects 21.4 26.2 27.5 39.6 34.0
    (4) No coherence for language templates 24.1 28.3 25.8 35.2 32.1
    (5) No relevance for language templates 29.5 34.0 32.8 40.4 29.4
    (6) No coherence (all components) 25.2 29.4 23.1 39.6 33.8
    (7) No relevance (all components) 9.6 11.9 11.0 36.8 23.9

    The data shows that completely ablating rationale relevance (setting 7) causes severe degradation, dropping GSM8K accuracy from 48.5% to 11.0% (worse than standard prompting at 15.4%). In contrast, providing invalid reasoning (setting 1) retains 39.5% accuracy on GSM8K and 39.4 F1 on Bamboogle. Additionally, ablating bridging object relevance (setting 3) causes greater performance decline than ablating bridging object coherence (setting 2), while ablating language template coherence (setting 4) severely damages performance.

  5. Knowl 5 — Qualitative Consistency and Invariance of Model Reasoning Under Invalid Demonstrations

    empirical result

    When large language models generate reasoning paths after being prompted with invalid demonstration rationales, the generated outputs remain overwhelmingly coherent, pertinent, and logically sound for correctly answered instances.

    Qualitative error analysis on GSM8K shows that the error distribution when prompting with invalid reasoning is nearly identical to standard CoT prompting:

    • Calculation errors: account for 20% of errors in both CoT and invalid reasoning configurations.
    • Missing intermediate steps: account for 35% of errors in CoT versus 25% under invalid reasoning.
    • Semantic understanding errors: account for 45% of errors in CoT versus 55% under invalid reasoning.

    The model generates valid deduction steps despite receiving demonstrations containing invalid leaps or incorrect arithmetic, indicating that in-context rationales function primarily as structural formatting templates rather than instructional logic rules.

  6. Knowl 6 — Primacy of Query Relevance and Language Template Coherence in Prompt Effectiveness

    empirical result

    Ablation experiments isolating the relevance and coherence of bridging objects versus language templates reveal asymmetric dependencies in CoT prompting:

    1. Relevance is Essential: Completely removing query relevance from both bridging objects and language templates (substituting numbers and entities with unrelated values from different domains) collapses model performance to levels at or below standard prompting without rationales (e.g., GSM8K intermediate F1 drops from 48.3% to 11.9% on text-davinci-002). Without query relevance, models frequently generate rationales centered on unrelated pretraining patterns.
    2. Relevance Over Coherence for Bridging Objects: For bridging objects, maintaining numerical/entity relevance is significantly more critical than preserving proper order (GSM8K intermediate F1 is 39.2% for incoherent bridging objects versus 26.2% for irrelevant bridging objects).
    3. Coherence is Critical for Language Templates: Scrambling the ordering of language templates causes substantial degradation (GSM8K intermediate F1 drops to 28.3%), inducing the model to generate incoherent natural language scaffolding during inference.
  7. Knowl 7 — Modulation of Demonstration Sensitivity by Model Scale and Task-Specific Instruction Tuning

    empirical result

    The impact of demonstration corruptions (invalid reasoning, loss of relevance, loss of coherence) decreases as the language model possesses stronger task-specific prior knowledge through scale or instruction tuning:

    • InstructGPT (text-davinci-003): Exhibits substantially higher robustness than text-davinci-002; prompting with invalid reasoning yields an intermediate F1 of 53.5% on GSM8K compared to 53.1% for valid CoT prompting.
    • Flan-PaLM: Having been explicitly instruction-tuned on arithmetic reasoning and factual QA formatted with chains of thought, Flan-PaLM achieves near-total invariance across all ablation settings. On GSM8K, standard CoT obtains 63.8% accuracy and 73.0% intermediate F1, while invalid reasoning yields 64.4% accuracy and 72.6% intermediate F1, and complete removal of relevance still yields 64.5% accuracy and 71.6% intermediate F1.

    This confirms that demonstrations in CoT prompting act largely to specify the output format and elicit existing pretraining priors rather than teaching task-specific inference in-context.

  8. Knowl 8 — Dual Intrinsic and Extrinsic Metric Evaluation Framework for Chain-of-Thought Rationale Quality

    experimental setup

    To evaluate Chain-of-Thought rationales beyond extrinsic final-answer correctness, an intrinsic evaluation framework measures the accuracy of generated intermediate reasoning paths:

    1. Extrinsic Metrics: Exact-match Answer Accuracy for arithmetic reasoning (GSM8K) and token-level Answer F1 for multi-hop QA (Bamboogle).
    2. Intrinsic Intermediate Metrics (Inter. Recall / Inter. F1): Computes the precision, recall, and F1 of derived intermediate bridging objects generated by the model. Derived bridging objects are defined as all numerical quantities (GSM8K) or entity mentions (Bamboogle) that appear in the ground-truth intermediate reasoning steps but do not appear in the original question prompt.

    Measuring intermediate bridging object overlap avoids assigning full credit to correct answers derived via flawed reasoning, and avoids assigning zero credit to valid multi-step reasoning chains that fail only on the final computation.

  9. Knowl 9 — Scope Limitations Regarding Template-Rigid Symbolic Reasoning and Formal Fallacy Modeling

    limitation

    The empirical findings on Chain-of-Thought prompting carry three primary limitations:

    1. Reasoning Task Scope: Experiments focus on multi-step arithmetic reasoning (GSM8K) and multi-hop factual question answering (Bamboogle). Symbolic reasoning benchmarks with rigid, invariant linguistic templates across examples (such as Last Letter Concatenation or Coin Flip) were excluded because their lack of step-level variation precludes permuting order or replacing templates independently.
    2. Informal Fallacy Synthesis: Invalid reasoning exemplars were manually written rather than generated through a formal, automated categorization of informal logical fallacies.
    3. Reference Dependency of Intrinsic Evaluation: Intrinsic intermediate metric evaluation relies on human-annotated ground truth intermediate entities and equations, which are resource-intensive to produce and not broadly available across arbitrary reasoning datasets.

Coverage note — None was omitted. All key findings, formal definitions of rationale components, ablation configurations, quantitative results across models, and stated limitations were fully extracted.

References

  1. 1.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2023. What learning algorithm is in-context learning? investigations with linear models. In The Eleventh International Conference on Learning Representations.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Ohio Supercomputer Center. 1987. Ohio supercomputer center.
  4. 4.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  5. 5.François Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  9. 9.Irving Copi, Carl Cohen, and Victor Rodych. 2016. Introduction to logic. Routledge.
  10. 10.Yao Fu, Hao Peng, and Tushar Khot. 2022. How does gpt obtain its ability? tracing emergent abilities of language models to their sources. Yao Fu’s Notion.
  11. 11.Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. 2022. What can transformers learn incontext? a case study of simple function classes. In Advances in Neural Information Processing Systems, volume 35, pages 30583–30598. Curran Associates, Inc.
  12. 12.Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. ROSCOE: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations.
  13. 13.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403.
  14. 14.Joel Jang, Seonghyeon Ye, and Minjoon Seo. 2023. Can large language models truly understand prompts? a case study with negated prompts. In Transfer Learning for Natural Language Processing Workshop, pages 52–62. PMLR.
  15. 15.Aman Madaan and Amir Yazdanbakhsh. 2022. Text and patterns: For effective chain of thought, it takes two to tango. arXiv preprint arXiv:2209.07686.
  16. 16.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  17. 17.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  18. 18.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  19. 19.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 840–854, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  20. 20.Abulhair Saparov and He He. 2023. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. In The Eleventh International Conference on Learning Representations.
  21. 21.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  22. 22.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  23. 23.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  24. 24.Albert Webson and Ellie Pavlick. 2022. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States. Association for Computational Linguistics.
  25. 25.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  26. 26.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846.
  27. 27.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations.
  28. 28.Xi Ye, Srinivasan Iyer, Asli Celikyilmaz, Ves Stoyanov, Greg Durrett, and Ramakanth Pasunuru. 2022. Complementary explanations for effective in-context learning. arXiv preprint arXiv:2211.13892.
  29. 29.Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang-goo Lee, and Taeuk Kim. 2022. Ground-truth labels matter: A deeper look into input-label demonstrations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2422–2437, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  30. 30.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Wang, B., et al. “Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 2717–39, https://doi.org/10.18653/v1/2023.acl-long.153.
APA
Wang, B., Min, S., Deng, X., Shen, J., Wu, Y., Zettlemoyer, L., & Sun, H. (2023). Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2717–2739. https://doi.org/10.18653/v1/2023.acl-long.153
Chicago
Wang, B., S. Min, X. Deng, et al. 2023. “Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2717–39. https://doi.org/10.18653/v1/2023.acl-long.153.
Harvard
Wang, B. et al. (2023) “Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2717–2739. Available at: https://doi.org/10.18653/v1/2023.acl-long.153.
Vancouver
1. Wang B, Min S, Deng X, Shen J, Wu Y, Zettlemoyer L, Sun H (2023) Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2717–2739

BibTeX

@inproceedings{wang-etal-2023-towards,
    title = "Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters",
    author = "Wang, Boshi  and
      Min, Sewon  and
      Deng, Xiang  and
      Shen, Jiaming  and
      Wu, You  and
      Zettlemoyer, Luke  and
      Sun, Huan",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.153/",
    doi = "10.18653/v1/2023.acl-long.153",
    pages = "2717--2739"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/