How Language Model Hallucinations Can Snowball

Muru ZhangOfir PressWilliam MerrillAlisa LiuNoah A. Smith

article2024ICML479 citations

Reveals that large language models frequently invent false justifications to maintain consistency with their own earlier mistakes, generating secondary errors that they are otherwise capable of correctly identifying as false in isolation.

Listen

As large language models are increasingly deployed in real-world information retrieval and decision-making systems, their tendency to produce plausible but false statements presents a major operational and safety risk. While conventional industry wisdom assumes these errors stem from gaps in training knowledge, the article demonstrates that language models frequently invent incorrect claims that they can separately recognize as false in isolation. This behavior, termed hallucination snowballing, occurs when a model commits to an incorrect initial response and subsequently manufactures false supporting details to maintain internal consistency.

The article evaluates why and how frequently this phenomenon occurs across leading commercial and open-source models, specifically testing GPT-3.5, GPT-4, and LLaMA2-70B-chat. The authors designed three targeted question-answering datasets of 500 questions each spanning distinct domains: primality testing, biographical senator verification, and multi-step flight connectivity. In a two-stage evaluation, models were first evaluated on whether they answered the queries correctly under standard zero-shot prompting. When an incorrect answer was produced, the supporting claims were extracted and fed back into the same model in separate sessions to verify whether the model possessed the correct underlying knowledge.

The findings reveal that models struggle heavily with single-step commitments on sequential reasoning tasks, exhibiting average error rates between 60% and 83%. More crucially, when models generated incorrect explanations to support their false answers, they were capable of identifying their own supporting claims as false 67.37% of the time for GPT-3.5, 87.03% for GPT-4, and 93.67% for LLaMA2-70B-chat. Standard decoding adjustments, including higher sampling temperatures and beam search, failed to alleviate this behavior. While zero-shot chain-of-thought prompting ("Let's think step by step") lowered error rates on simple benchmarks for GPT-3.5 and GPT-4, it proved ineffective on composite, multi-part reasoning questions, where both models still failed on more than half of the questions and continued to experience snowballed errors in over 65% of those failures.

These results indicate that popular deployment mitigations, such as knowledge retrieval augmentation, are insufficient to solve hallucinations because the root problem is structural rather than purely informational. Models prioritize output consistency over factuality once an initial token commitment is made. For organizations utilizing language models, this creates hidden compliance and performance risks, as confident justifications cannot be treated as evidence of factual reasoning.

To address this systemic risk, the article suggests re-evaluating model training and generation workflows. System developers should implement architectures that force explicit reasoning chains prior to answering, and explore fine-tuning techniques that teach models how to backtrack and revise prior statements rather than forcing consistency. In the interim, decision-makers should recognize that model self-confidence is fragile, and critical applications should avoid relying on single-pass model outputs without independent, modular verification pipelines.

Cover for How Language Model Hallucinations Can Snowball

Abstract

A major risk of using language models in practical applications is their tendency to hallucinate incorrect statements. Hallucinations are often attributed to knowledge gaps in LMs, but we show that LMs sometimes produce hallucinations that they can separately recognize as incorrect. To do this, we construct three question-answering datasets where LMs often state an incorrect answer which is followed by an explanation with at least one incorrect claim. Crucially, we find that GPT-3.5, GPT-4, and LLaMA2-70B-chat can identify 67%, 87%, and 94% of these incorrect claims, respectively. We show that this phenomenon doesn’t disappear under higher temperatures sampling, beam search, and zero-shot chain-of-thought prompting. These findings reveal that LM hallucinations can snowball: early mistakes by an LM can lead to more mistakes that otherwise would not be made.

Table of Contents

  • 1. Introduction
  • 2. Why Do We Expect Hallucination Snowballing?
  • 3. Experiments
  • 3.1. Datasets
  • 3.2. Inference Setup
  • 3.3. LM Recognition of Snowballed Hallucinations
  • 3.4. Results
  • 4. Zero-shot Chain-of-thought Prompting
  • 4.1. “Let’s think step by step” can Alleviate Hallucination
  • 4.2. Does “Let’s think step by step” Work on Composite Questions?
  • 4.3. “Let’s think step by step” Struggles on the Composite Datasets
  • 5. Towards a Robust Solution for Snowballing Hallucinations
  • 6. Related Work
  • 7. Conclusion
  • Impact Statement
  • References
  • A. Dataset Details
  • A.1. Graph Connectivity
  • A.2. Senator search
  • B. Additional Results

Knowls

  1. Knowl 1 — Definition of Hallucination Snowballing

    definition

    Hallucination snowballing is the phenomenon in which an autoregressive language model generates one or more factually incorrect claims within an explanation solely to maintain consistency with an earlier incorrect assertion or answer (such as an immediate "Yes" or "No" committal), even though the model separately recognizes those explanatory claims as incorrect when queried about them in isolation.

    Snowballed hallucinations are distinguished from standard hallucinations caused by "knowledge gaps":

    • Knowledge-gap hallucination: The model generates false information because the underlying facts were absent from or incorrectly learned during pretraining.
    • Snowballed hallucination: The model possesses the requisite factual knowledge (and verifies it correctly when queried independently in a clean context), but fabricates false premises or false deductions to justify a prior erroneous generation.
  2. Knowl 2 — Theoretical Mechanism of Hallucination Snowballing via Initial Committal and Sequential Complexity

    theoretical result

    Language models are systematically vulnerable to hallucination snowballing on reasoning tasks due to the interaction of three structural factors:

    1. Initial Committal: Instruction-tuned autoregressive language models exhibit an answer-first bias on yes/no questions, outputting the categorical decision token before generating reasoning. Empirically, the first output token is "Yes" or "No" in 95.67%95.67\% of completions for GPT-4 and 98.40%98.40\% for GPT-3.5 under direct zero-shot prompting.

    2. Single-Step Computational Inability for Inherently Sequential Tasks: Bounded-precision transformer architectures cannot solve computational problems outside the circuit complexity class TC0\mathsf{TC}^0 within a single generation step. Inherently sequential problems—such as directed graph connectivity (which is L\mathsf{L}-complete) and primality testing (in P\mathsf{P}, but outside TC0\mathsf{TC}^0 unless P⊆L\mathsf{P} \subseteq \mathsf{L})—require multi-step computation. Forcing the model to commit to an answer token at timestep 1 forces it to make an uncalculated guess on problems beyond its single-step capacity.

    3. Conditioned Context Consistency over Truthfulness: Once an incorrect answer token (e.g., asserting that a prime number is composite) is generated, it is appended to the autoregressive context. During subsequent token generation, conditional probability favors discourse coherence and justification of the prior token over external factuality. The model therefore generates fabricated supporting evidence (e.g., fictitious factorizations or nonexistent edges) that it would reject in an unconditioned context.

  3. Knowl 3 — Two-Stage Verification Protocol for Identifying Snowballed Hallucinations

    model/method

    To determine whether an erroneous explanation constitutes a snowballed hallucination rather than a factual knowledge deficit, a two-stage evaluation protocol is employed:

    1. Task Execution and Claim Extraction:

      • A language model is queried on a task designed such that the ground truth is fixed across all instances (e.g., all numbers are prime, or no path exists in a disconnected graph).
      • If the model provides an incorrect initial answer, it generates an explanation justifying that answer.
      • Specific factual claims within the generated explanation are extracted: candidate integer factors aa and bb for primality questions, a directed edge (u,v)(u, v) for flight connectivity paths, or a politician's claimed biographical relation (S,X,Y)(S, X, Y) for senator queries.
    2. Isolated Verification Probing:

      • In an entirely new interaction session (with no prior conversation history or context), the exact same model is presented with the extracted claim in isolation (e.g., "Is NN divisible by aa? Answer with either Yes or No.", or "Based on the above flight information, is City uu to City vv a valid flight?").
      • If the model correctly refutes the claim in isolation (e.g., answering that NN is not divisible by aa), the original explanatory claim is classified as a snowballed hallucination.
  4. Knowl 4 — Benchmark QA Datasets for Probing Single-Step Hallucination Snowballing

    experimental setup

    Three question-answering (QA) datasets of 500 yes/no questions each were constructed to evaluate single-step committal and hallucination snowballing, where ground-truth answers are constant to facilitate deterministic verification:

    1. Primality Testing (500 instances): Prompts query whether a randomly sampled integer N∈[1000,20000]N \in [1000, 20000] is prime. All selected integers are prime (ground truth: "Yes"). Incorrect "No" answers require the model to propose integer factors a×b=Na \times b = N.

    2. Senator Search (500 instances): Prompts query whether a U.S. senator ever represented state XX and graduated from university YY, drawn from all 50 U.S. states and a list of 12 selective universities: MIT, University of Chicago, Johns Hopkins University, California Institute of Technology, Duke University, Northwestern University, Dartmouth College, Brown University, Vanderbilt University, Rice University, and University of Washington (excluding universities on U.S. News Top 10 Colleges for Members of Congress). Pairs where a senator actually existed were removed (ground truth: "No"). Incorrect "Yes" answers require naming a specific senator.

    3. Graph Connectivity (500 instances): Prompts present a 12-flight directed graph among 14 cities with randomly assigned English letter names. The underlying structure consists of two disconnected subgraphs. Queries ask whether a valid path connects source city ss (a source node in subgraph 1) to destination city tt (a leaf node in subgraph 2), making 1-step heuristics inapplicable (ground truth: "No"). Incorrect "Yes" answers produce candidate paths containing nonexistent edges.

  5. Knowl 5 — Snowballed Hallucination Rates Under Direct Prompting

    empirical result

    Across greedy decoding evaluations on GPT-3.5 (gpt-3.5-turbo), GPT-4, and LLaMA-2-70B-chat, models achieve low overall accuracy under direct prompting, but successfully recognize the vast majority of their explanatory errors when queried in isolation.

    Model Graph Connectivity Primality Testing Senator Search Average
    Task Error Rate (Mistakes / 500)
    GPT-3.5 410/500 (82.0%) 339/500 (67.8%) 153/500 (30.6%) 60.13%
    GPT-4 442/500 (88.4%) 374/500 (74.8%) 435/500 (87.0%) 83.40%
    LLaMA-2-70B-chat 487/500 (97.4%) 248/500 (49.6%) 265/500 (53.0%) 66.67%
    Snowballed Hallucinations / Total Errors
    GPT-3.5 396/410 (96.6%) 125/339 (36.9%) 98/153 (68.6%) 67.37%
    GPT-4 417/442 (94.3%) 346/374 (92.5%) 323/435 (74.3%) 87.03%
    LLaMA-2-70B-chat 474/487 (97.3%) 215/248 (86.7%) 257/265 (97.0%) 93.67%

    These results establish that the majority of hallucinations in multi-step QA are not caused by missing knowledge, but are generated downstream of an initial incorrect committal to maintain coherence.

  6. Knowl 6 — Impact of Zero-Shot Chain-of-Thought Prompting on Simple Reasoning Benchmarks

    empirical result

    Appending the zero-shot Chain-of-Thought (CoT) trigger "Let's think step by step" substantially reduces initial committal and task error rates for GPT-3.5 and GPT-4 on simple single-problem tasks, but fails to improve LLaMA-2-70B-chat. However, when errors do occur during the generation of the reasoning chain, hallucination snowballing persists at very high rates.

    Model Graph Connectivity Primality Testing Senator Search Average
    Error Rate with "Let's think step by step"
    GPT-3.5 139/500 (27.8%) 2/500 (0.4%) 0/500 (0.0%) 9.40%
    GPT-4 21/500 (4.2%) 37/500 (7.4%) 0/500 (0.0%) 3.87%
    LLaMA-2-70B-chat 436/500 (87.2%) 367/500 (73.4%) 200/500 (40.0%) 66.87%
    Snowballed Hallucinations / Total Errors
    GPT-3.5 123/139 (88.5%) 0/2 (0.0%) 0/0 (N/A) 44.25%
    GPT-4 20/21 (95.2%) 35/37 (94.6%) 0/0 (N/A) 94.90%
    LLaMA-2-70B-chat 427/436 (97.9%) 313/367 (85.3%) 197/200 (98.5%) 93.90%

    While CoT delays initial committal on straightforward prompts, errors introduced in early reasoning steps continue to cause snowballed hallucinations in downstream steps.

  7. Knowl 7 — Composite QA Benchmarks for Multi-Problem Concurrent Reasoning

    experimental setup

    To evaluate whether zero-shot Chain-of-Thought prompting remains effective when text generation involves concurrent reasoning over multiple sub-problems, three composite datasets containing 500 questions each were constructed:

    1. Composite Primality Testing: Prompts present three randomly selected prime numbers p1,p2,p3∈[1000,20000]p_1, p_2, p_3 \in [1000, 20000] in the format: "Are any of the following numbers not prime: p1, p2, p3?". Ground truth is always "No".

    2. Composite Senator Search: Prompts query whether any of three candidate U.S. states had a senator who attended a designated university: "Which of the following states have ever had senators who went to [University]: State 1, State 2, or State 3?". No state in the tuple satisfies the condition, and no three-state tuple is reused (ground truth: "None").

    3. Composite Graph Connectivity: Prompts present the 12-flight network and ask the model to evaluate three separate source-target pairs simultaneously: "For each of the following pairs of cities, check if there is a series of flights from the source city to the target city: (s1, t1), (s2, t2), (s3, t3)". All three paths are invalid (ground truth: "No" for all three).

  8. Knowl 8 — Failure of Zero-Shot Chain-of-Thought on Composite Reasoning Datasets

    empirical result

    When evaluated on composite datasets requiring concurrent evaluation of three sub-problems, zero-shot Chain-of-Thought prompting ("Let's think step by step") fails to prevent initial committal and hallucination snowballing, leading to high error rates and high proportions of self-refutable claims.

    Model Composite Graph Composite Primality Composite Senator Average
    Error Rate with "Let's think step by step"
    GPT-3.5 491/500 (98.2%) 214/500 (42.8%) 198/500 (39.6%) 60.20%
    GPT-4 200/500 (40.0%) 352/500 (70.4%) 301/500 (60.2%) 56.87%
    Snowballed Hallucinations / Total Errors
    GPT-3.5 485/491 (98.8%) 86/214 (40.2%) 120/198 (60.6%) 66.53%
    GPT-4 171/200 (85.5%) 279/352 (79.3%) 206/301 (68.4%) 77.73%

    Both models hallucinate on more than half of the composite instances, and over 65%65\% of these hallucinations snowball. The models commit early to a partial claim about one sub-element, forcing subsequent erroneous justifications.

  9. Knowl 9 — Robustness of Hallucination Snowballing Across Decoding Strategies

    empirical result

    Modifying inference-time decoding parameters—including sampling temperature, top-kk, nucleus sampling, and beam search—does not eliminate hallucination snowballing.

    • Temperature Scaling (t∈{0.0,0.6,0.9}t \in \{0.0, 0.6, 0.9\}): Increasing temperature spreads probability mass away from the mode but leaves error rates and snowballing rates virtually unchanged across all models:
      • GPT-3.5 average error rate: 60.13%60.13\% (t=0.0t=0.0), 58.53%58.53\% (t=0.6t=0.6), 58.53%58.53\% (t=0.9t=0.9); snowballing proportions remain 67.37%67.37\%, 66.77%66.77\%, and 66.77%66.77\%.
      • GPT-4 average error rate: 83.40%83.40\% (t=0.0t=0.0), 82.53%82.53\% (t=0.6t=0.6), 81.67%81.67\% (t=0.9t=0.9); snowballing proportions remain 87.03%87.03\%, 86.13%86.13\%, and 84.87%84.87\%.
      • LLaMA-2-70B-chat average error rate: 66.67%66.67\% (t=0.0t=0.0), 70.67%70.67\% (t=0.6t=0.6), 67.33%67.33\% (t=0.9t=0.9); snowballing proportions remain 93.67%93.67\%, 95.07%95.07\%, and 94.07%94.07\%.
    • Top-kk and Nucleus Sampling: Truncating lower-probability tails restricts sampling to the top tokens, which only reinforces the probability of immediate answer committal.
    • Beam Search: Beam search with width k=10k=10 evaluated on LLaMA-2-70B-chat reduced error on Primality Testing from 49.6%49.6\% to 35.2%35.2\%, but increased error on Senator Search from 53.0%53.0\% to 73.0%73.0\%, leaving average error unchanged at 66.67%66.67\% and resulting in an 89.93%89.93\% snowballed hallucination rate among errors.
  10. Knowl 10 — Exposure Bias and Absence of Backtracking in Autoregressive Training

    limitation

    Hallucination snowballing is fundamentally linked to training-distribution biases and the lack of backtracking mechanisms in current large language models:

    1. Exposure Bias: Standard pretraining uses teacher forcing over gold-standard corpora, conditioning generation strictly on true prior prefixes. During inference, models must condition on their own previously generated tokens, including unverified or erroneous commitments.
    2. Absence of Revision Patterns: Natural human writing and web-scraped text corpora typically contain finalized text where initial drafting errors and internal self-corrections have been scrubbed. Consequently, models do not learn explicit backtracking behaviors (e.g., generating "Sorry, that was incorrect" and amending a prior assertion). Instead, the model's conditional language modeling objective forces it to treat all prior context as ground truth, compelling it to double down on earlier missteps.

Coverage note — No substantial contributed material was omitted. All task setups, datasets (base and composite), theoretical complexity motivations, verification protocols, prompting variations (direct vs. CoT), decoding ablations (temperature, beam search), and underlying causes are covered.

References

  1. 1.Agrawal, M., Kayal, N., and Saxena, N. Primes is in p. Annals of Mathematics, 160:781–793, June 2004. URL https://www.microsoft.com/en-us/research/publication/primes-is-in-p/. Godel Prize, Fulkerson Prize.
  2. 2.Arora, K., El Asri, L., Bahuleyan, H., and Cheung, J. Why exposure bias matters: An imitation learning perspective of error accumulation in language generation. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Findings of the Association for Computational Linguistics: ACL 2022, pp. 700–710, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022. findings-acl.58. URL https://aclanthology.org/2022.findings-acl.58.
  3. 3.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  4. 4.Kim, G., Baldi, P., and McAleer, S. Language models can solve computer tasks, 2023.
  5. 5.Kim, N., Pavlick, E., Ayan, B. K., and Ramachandran, D. Which linguist invented the lightbulb? presupposition verification for question-answering. In Annual Meeting of the Association for Computational Linguistics, 2021.
  6. 6.Kim, N., Htut, P. M., Bowman, S., and Petty, J. (qa)2: Question answering with questionable assumptions. ArXiv, abs/2212.10003, 2022.
  7. 7.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners, 2023.
  8. 8.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS'20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  9. 9.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.
  10. 10.Liu, N. F., Zhang, T., and Liang, P. Evaluating verifiability in generative search engines, 2023.
  11. 11.Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. On faithfulness and factuality in abstractive summarization. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 1906–1919, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.173. URL https://aclanthology.org/2020.acl-main.173.
  12. 12.Merrill, W. and Sabharwal, A. The parallelism tradeoff: Limitations of log-precision transformers, 2023.
  13. 13.Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A. Show your work: Scratchpads for intermediate computation with language models, 2021.
  14. 14.OpenAI. Introducing chatgpt. 2022. URL https://openai.com/blog/chatgpt.
  15. 15.OpenAI. Gpt-4 technical report, 2023.
  16. 16.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022.
  17. 17.Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., Chen, W., and Gao, J. Check your facts and try again: Improving large language models with external knowledge and automated feedback, 2023.
  18. 18.Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N. A., and Lewis, M. Measuring and narrowing the compositionality gap in language models, 2022.
  19. 19.Raunak, V., Menezes, A., and Junczys-Dowmunt, M. The curious case of hallucinations in neural machine translation. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1172–1183, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.92. URL https://aclanthology.org/2021.naacl-main.92.
  20. 20.Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K. Object hallucination in image captioning. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4035–4045, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1437. URL https://aclanthology.org/D18-1437.
  21. 21.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Bers, T., Biderman, S., Gao, L., Wolf, T., and Rush, A. M. Multitask prompted training enables zero-shot task generalization, 2021.
  22. 22.Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J. Retrieval augmentation reduces hallucination in conversation. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3784–3803, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp. 320. URL https://aclanthology.org/2021.findings-emnlp.320.
  23. 23.Wang, C. and Sennrich, R. On exposure bias, hallucination and domain shift in neural machine translation. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 3544–3552, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.326. URL https://aclanthology.org/2020.acl-main.326.
  24. 24.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions, 2022.
  25. 25.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners, 2021.
  26. 26.Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=_VjQlMeSB_J.
  27. 27.Yu, X. V., Min, S., Zettlemoyer, L., and Hajishirzi, H. Crepe: Open-domain question answering with false presuppositions, 2022.
  28. 28.Zheng, S., Huang, J., and Chang, K. C.-C. Why does chatgpt fall short in answering questions faithfully?, 2023.

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/