The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

Aviv SlobodkinOmer GoldmanAvi CaciularuIdo DaganShauli Ravfogel

article2023EMNLP74 citations

Reveals that large language models internally encode whether a question is answerable in their hidden states even while generating hallucinatory answers, showing that this linearly separable answerability signal can be extracted to curb overconfident errors.

Listen

Large language models frequently generate confident but incorrect responses when faced with unanswerable questions. In real-world applications such as search engines and automated assistants, these fabricated answers create substantial compliance, safety, and reliability risks. The article investigates whether language models inherently track the answerability of a question within their internal representations even while generating incorrect text, and whether this internal state can be used to prevent such errors.

The researchers evaluated three instruction-tuned models across three reading comprehension benchmarks containing both answerable and unanswerable questions. They applied four distinct analytical techniques: modifying prompts to explicitly give models permission to abstain, inspecting alternative potential answers during text generation, training simple linear classifiers on the models' internal hidden layers to predict answerability, and mathematically erasing the identified answerability features to verify their direct influence on model behavior.

The findings show that language models internally recognize when a question cannot be answered from the provided text. First, simply including instructions that explicitly permit abstention improved unanswerability detection performance by up to 80 points, with overall question-answering accuracy increasing by more than 50 points. Second, when standard models fabricated an answer, correct abstention responses were typically present as lower-probability candidates in the generated beam. Third, a basic linear classifier trained on the internal state of the first generated word achieved over 75% accuracy across all models and benchmarks, demonstrating that answerability is clearly organized in the model's internal data space. Finally, erasing this internal feature substantially degraded performance, confirming its functional importance.

These results indicate that overconfident inaccuracies are not caused by a failure of model comprehension, but rather by generation mechanisms that prioritize generating a definitive answer over admitting ignorance. Organizations deploying language models should immediately incorporate explicit abstention instructions into system prompts. Furthermore, engineering teams can implement lightweight internal classifiers or search candidate response sets to filter out unanswerable queries before responses reach end users.

The primary limitations include a focus on reading comprehension tasks within specific text contexts rather than open-domain knowledge queries, as well as testing a selected set of model architectures. Future work should assess how these internal mechanisms function in open-domain environments and evaluate larger model sizes.

Cover for The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

Abstract

Large language models (LLMs) have been shown to possess impressive capabilities, while also raising crucial concerns about the faithfulness of their responses. A primary issue arising in this context is the management of (un)answerable queries by LLMs, which often results in hallucinatory behavior due to overconfidence. In this paper, we explore the behavior of LLMs when presented with (un)answerable queries. We ask: do models represent the fact that the question is (un)answerable when generating a hallucinatory answer? Our results show strong indications that such models encode the answerability of an input query, with the representation of the first decoded token often being a strong indicator. These findings shed new light on the spatial organization within the latent representations of LLMs, unveiling previously unexplored facets of these models. Moreover, they pave the way for the development of improved decoding techniques with better adherence to factual generation, particularly in scenarios where query (un)answerability is a concern.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Prompt Manipulation
  • 3.2 Beam Relaxation
  • 3.3 Identifying an Answerability Subspace
  • 3.4 Erasing the Answerability Subspace
  • 4 Experimental Setup
  • 5 Results
  • 5.1 Prompt Manipulation
  • 5.1.1 Zero-Shot Scenario
  • 5.1.2 Few-Shot Scenario
  • 5.2 Beam Relaxation
  • 5.3 Identifying an Answerability Subspace
  • 5.4 Erasing the Answerability Subspace
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics Statement
  • References
  • A (Un)answerability-Recognizing Responses
  • B Performance on the Answerable Instances
  • C Prompt Variant Tuning
  • D Impact of Hint Placement
  • E Impact of the Relaxed Beam Search Decoding on the Answerable Queries
  • F In-Context-Learning Variants

Knowls

  1. Knowl 1 — Answerability is linearly decodable from first-token hidden states

    empirical result

    For three instruction-tuned language models—Flan-T5xxl (11B parameters), Flan-UL2 (20B), and OPT-IML (30B)—a logistic-regression classifier can distinguish answerable from unanswerable questions using the last-layer representation of the first generated token. The classifier was trained and evaluated separately for each model, benchmark, and prompt condition; unanswerability was the positive class. The comparison feature was the initial, non-contextual embedding, which should contain no question-specific answerability information. The table reports F1 scores (percent):

    ModelPromptSQuAD initialSQuAD lastNQ initialNQ lastMuSiQue initialMuSiQue last
    Flan-T5xxlRegular40.189.923.086.147.477.5
    Flan-T5xxlHint40.089.426.686.238.677.3
    Flan-UL2Regular39.490.442.287.315.178.3
    Flan-UL2Hint39.689.941.587.941.678.3
    OPT-IMLRegular48.482.840.885.545.675.5
    OPT-IMLHint48.483.945.386.245.684.9

    The contextual last-layer features score 75.5–90.4 F1, whereas the initial-embedding baseline scores 15.1–48.4 F1. The strong performance with regular prompts shows that answerability information is present even when the prompt does not explicitly invite abstention; adding a hint changes the probe scores only modestly.

  2. Knowl 2 — An abstention hint sharply improves answerability classification

    empirical result

    The models were evaluated on whether they abstained on unanswerable questions, with abstention counted as an unanswerability prediction and any attempted answer counted as an answerability prediction, irrespective of answer correctness. The metric is F1 with unanswerable as the positive class. A hint explicitly allowed abstention; in few-shot prompts, the regular version used two answerable examples, while the hint version included one unanswerable example. Few-shot values are means across three example variants, with standard deviations in parentheses.

    ModelSettingPromptSQuAD F1NQ F1MuSiQue F1
    Flan-T5xxl (11B)Zero-shotRegular37.35.78.2
    Flan-T5xxl (11B)Zero-shotHint91.585.674.7
    Flan-UL2 (20B)Zero-shotRegular46.313.86.8
    Flan-UL2 (20B)Zero-shotHint92.383.967.3
    OPT-IML (30B)Zero-shotRegular43.921.817.1
    OPT-IML (30B)Zero-shotHint85.485.354.1
    Flan-T5xxl (11B)Few-shotRegular52.6 (11.7)8.6 (2.1)13.6 (4.9)
    Flan-T5xxl (11B)Few-shotHint91.2 (1.0)85.8 (0.5)75.4 (1.4)
    Flan-UL2 (20B)Few-shotRegular67.7 (4.0)20.6 (1.8)14.5 (4.9)
    Flan-UL2 (20B)Few-shotHint92.5 (0.1)83.7 (0.1)72.1 (0.6)
    OPT-IML (30B)Few-shotRegular10.0 (1.2)11.1 (2.1)4.9 (1.0)
    OPT-IML (30B)Few-shotHint79.3 (0.3)85.0 (1.5)27.5 (8.0)

    The hint raises F1 in every model–benchmark comparison in both settings. Gains are especially large on NQ, including 5.7 to 85.6 F1 for zero-shot Flan-T5xxl and 11.1 to 85.0 for few-shot OPT-IML. Thus, merely making abstention an available response substantially changes how often the models recognize unanswerability.

  3. Knowl 3 — The hint also improves end-to-end question answering

    empirical result

    On the full question-answering evaluation, models received credit for correctly answering answerable questions and abstaining on unanswerable ones. Exact match (EM) and token-level F1 are reported for zero-shot prompting below; each cell pair is EM/F1. A hint improves both metrics for every model on all three benchmarks.

    ModelPromptSQuAD EM/F1NQ EM/F1MuSiQue EM/F1
    Flan-T5xxl (11B)Regular55.4 / 58.622.5 / 27.041.2 / 48.5
    Flan-T5xxl (11B)Hint86.5 / 89.673.3 / 77.063.5 / 70.5
    Flan-UL2 (20B)Regular59.5 / 62.321.7 / 27.335.3 / 43.4
    Flan-UL2 (20B)Hint87.8 / 90.570.5 / 74.555.3 / 62.2
    OPT-IML (30B)Regular57.8 / 60.629.0 / 33.037.3 / 44.4
    OPT-IML (30B)Hint81.2 / 83.573.9 / 76.747.7 / 54.5

    The largest displayed F1 increase is on NQ for Flan-T5xxl, from 27.0 to 77.0; the largest EM increase is also on NQ, for OPT-IML, from 29.0 to 73.9. The hint slightly reduces performance on answerable-only examples on average—by 8.3% in F1 and 7.1% in EM in the zero-shot setting—but the gain from treating unanswerable queries appropriately produces higher overall QA scores. The authors also report that the same overall improvement pattern holds in few-shot evaluation.

  4. Knowl 4 — Answerability probes transfer across QA datasets

    empirical result

    A linear answerability classifier trained on hidden-state representations from one benchmark can classify examples from another benchmark without retraining. The values below are F1 scores (percent); columns identify the classifier's training dataset and rows identify its test dataset. Each model has its own matrix.

    ModelTest datasetTrained on SQuADTrained on NQTrained on MuSiQue
    Flan-T5xxlSQuAD89.974.975.0
    Flan-T5xxlNQ84.186.176.1
    Flan-T5xxlMuSiQue71.565.577.5
    Flan-UL2SQuAD90.485.664.7
    Flan-UL2NQ82.487.382.5
    Flan-UL2MuSiQue73.473.578.3
    OPT-IMLSQuAD82.878.017.4
    OPT-IMLNQ84.485.541.5
    OPT-IMLMuSiQue63.660.075.5

    Transfer scores are generally above the initial-embedding baselines, supporting the interpretation that the probes detect answerability information beyond dataset-specific surface features. Transfer is not uniform: for OPT-IML, a probe trained on MuSiQue obtains only 17.4 F1 on SQuAD and 41.5 on NQ, whereas most other cross-dataset combinations remain substantially stronger.

  5. Knowl 5 — Beam relaxation promotes abstention when an abstaining response is in the beam

    algorithm

    Beam relaxation is a zero-shot decoding intervention that changes the selected answer only when the final beam contains a response recognized as abstaining. The method was evaluated with beam sizes k=1,3,5,7k=1,3,5,7.

    Input: Question and context, a language model, beam size k
    Output: A decoded response, possibly replaced with an abstention response
    1. Generate the final k candidate responses using beam search.
    2. Search the candidates for an abstention response. Recognized forms include
       Unanswerable, N/A, I don't know, IDK, Not known, Answer not in context,
       Unknown, No answer, It is unknown, None of the above, and The answer is unknown,
       including lowercase variants.
    3. If any candidate matches a recognized abstention response, return Unanswerable.
    4. Otherwise, return the highest-probability beam response.

    Across the tested models and benchmarks, increasing beam size consistently increased the chance that an abstention-recognizing response appeared somewhere in the beam, even though the top-beam answer under ordinary decoding changed little. Consequently, answerability-classification F1 improved with beam relaxation. Its effect on answerable questions was limited: across the tested settings, the reported maximum reductions relative to regular beam search were 7.7 points in EM and 8.7 points in F1.

  6. Knowl 6 — Erasing the answerability direction degrades model behavior

    empirical result

    To test whether answerability representations influence behavior rather than merely encode information, the authors applied closed-form linear concept erasure to Flan-UL2's first-generation-step representations. The erasure projection was fit using regular-prompt SQuAD training representations and applied during inference on the same model–dataset pairing. The answerability-classification F1 fell from 50.1 to 31.2 with regular beam decoding and from 65.4 to 32.7 with beam relaxation.

    On zero-shot SQuAD with beam size 3, QA metrics also declined after erasure. Each cell below is EM/F1; the answerable-only columns evaluate just answerable questions.

    Decoding methodErasureAll questions EM/F1Answerable questions EM/F1
    Regular beamNo60.2 / 63.887.1 / 94.1
    Regular beamYes50.2 / 55.182.0 / 91.8
    Relaxed beamNo60.9 / 64.487.1 / 94.1
    Relaxed beamYes50.8 / 55.782.0 / 91.8

    The decline in answerability classification and overall QA after removing the linearly decodable information supports the authors' claim that this subspace is behaviorally relevant. Erasure also lowers answerable-only scores, so the intervention is not selective to unanswerable examples.

  7. Knowl 7 — Hallucinated unanswerable cases occupy an intermediate region in latent-space projections

    empirical result

    Three-dimensional PCA visualizations of the first generated token's last-layer representations show distinct patterns for answerable examples, unanswerable examples correctly recognized as unanswerable, and unanswerable examples that elicited an answer. The visualizations cover Flan-T5xxl, Flan-UL2, and OPT-IML on SQuAD, NQ, and MuSiQue, with regular and hint prompts. Correctly recognized answerable and unanswerable instances form visibly separated groups, particularly with a hint in the prompt. Unanswerable instances that the model answered instead appear in a separate region between the answerable group and the correctly recognized unanswerable group. This intermediate placement is consistent with the authors' interpretation that the representations reflect uncertainty even when the generated response is overconfident; the PCA plots are qualitative evidence rather than a quantitative measure of uncertainty.

  8. Knowl 8 — Evaluation uses constructed contextual QA test sets and three instruction-tuned models

    experimental setup

    The experiments evaluate Flan-T5xxl (11B parameters), Flan-UL2 (20B), and OPT-IML (30B) on contextual question answering using the development sets of SQuAD 2.0, Natural Questions (NQ), and MuSiQue. The resulting test-set counts are:

    BenchmarkAnswerable examplesUnanswerable examples
    SQuAD59285945
    NQ34897719
    MuSiQue19501316

    For answerable NQ examples, the selected long answer is the context and the short answer is the target response. For unanswerable NQ examples, the question is paired with the paragraph from its source passage that has the highest Sentence-BERT cosine similarity to the question, excluding the annotated long answer. For answerable MuSiQue examples, the paragraphs aligned with the single-hop subquestions are concatenated as context. For unanswerable MuSiQue examples, a semantically closest paragraph is selected for unanswered subquestions and combined with the paragraphs for the remaining subquestions. SQuAD examples use their associated paragraph.

    For probing, the authors sampled 1000 training examples per benchmark, balanced between the two answerability classes; 800 examples were used to train a probe and 200 for probe development. The main answerability metric treats unanswerable as positive and counts any attempted answer as answerable regardless of its correctness. QA quality is separately measured by exact match and token-level F1.

  9. Knowl 9 — Instruction hints matter more than hints shown only in examples

    empirical result

    In few-shot prompting, the authors compared a hint in both the instructions and exemplars, in the instructions alone, and in the exemplars alone. Each score is answerability-classification F1, averaged across three in-context example variants; standard deviations are in parentheses. Instruction-only prompting nearly matches the combined hint, while exemplar-only prompting is substantially weaker, particularly on MuSiQue and for OPT-IML.

    ModelHint placementSQuAD F1NQ F1MuSiQue F1
    Flan-T5xxlInstructions and exemplars91.2 (1.0)85.8 (0.5)75.4 (1.4)
    Flan-T5xxlInstructions only91.1 (1.0)85.0 (0.2)74.2 (1.3)
    Flan-T5xxlExemplars only88.7 (2.1)81.9 (0.8)51.3 (7.2)
    Flan-UL2Instructions and exemplars92.5 (0.1)83.7 (0.1)72.1 (0.6)
    Flan-UL2Instructions only92.3 (0.1)82.3 (0.1)71.3 (1.0)
    Flan-UL2Exemplars only92.0 (0.1)79.7 (0.6)59.0 (1.7)
    OPT-IMLInstructions and exemplars79.3 (0.3)85.0 (1.5)27.5 (8.0)
    OPT-IMLInstructions only75.0 (1.8)81.9 (1.1)14.1 (4.2)
    OPT-IMLExemplars only58.1 (1.4)61.0 (6.9)8.4 (3.7)

    For Flan-T5xxl and Flan-UL2, adding the hint to the examples after it is already present in the instructions yields only small additional changes. OPT-IML benefits more from the combined condition than from instructions alone, but its exemplar-only results remain the weakest.

  10. Knowl 10 — The evidence is limited to selected models, benchmarks, and linear methods

    limitation

    The experiments cover three instruction-tuned models and three contextual QA benchmarks, so they do not establish how answerability encoding varies across model families or scales. The probing and erasure analyses use linear methods; because language-model representations may encode answerability nonlinearly, the reported linear results are a lower bound on what may be recoverable. The experiments do not compare the prompt, beam-relaxation, and representation-intervention approaches against one another, and they address answerability given a context rather than open-domain QA. The findings therefore provide evidence about these settings, not a general solution to hallucination.

Coverage note — Detailed few-shot example texts and the development-set comparison of alternative abstention labels were omitted because they serve as prompt illustrations or tuning details rather than distinct findings; the hint-placement ablation and the paper's principal limitations are included.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 298–306.
  2. 2.Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017. Fine-grained analysis of sentence embeddings using auxiliary prediction tasks. In International Conference on Learning Representations (ICLR).
  3. 3.Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang. 2023. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models.
  4. 4.Galen Andrew and Jianfeng Gao. 2007. Scalable training of L1-regularized log-linear models. In Proceedings of the 24th International Conference on Machine Learning, pages 33–40.
  5. 5.Akari Asai and Eunsol Choi. 2021. Challenges in information-seeking QA: Unanswerable questions and paragraph retrieval. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1492–1504, Online. Association for Computational Linguistics.
  6. 6.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. Leace: Perfect linear concept erasure in closed form. arXiv preprint arXiv:2306.03819.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  8. 8.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations.
  9. 9.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  10. 10.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  11. 11.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia. Association for Computational Linguistics.
  12. 12.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  13. 13.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models.
  14. 14.Shehzaad Dhuliawala, Leonard Adolphs, Rajarshi Das, and Mrinmaya Sachan. 2022. Calibration of machine reading systems at scale. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1682–1693, Dublin, Ireland. Association for Computational Linguistics.
  15. 15.Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2023. Jump to conclusions: Shortcutting transformers with linear transformations. arXiv preprint arXiv:2303.09435.
  16. 16.Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175.
  17. 17.James Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot, and Pradeep Dasigi. 2020. IIRC: A dataset of incomplete information reading comprehension questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1137–1147, Online. Association for Computational Linguistics.
  18. 18.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767.
  19. 19.Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg. 2022a. LM-debugger: An interactive tool for inspection and intervention in transformer-based language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 12–21, Abu Dhabi, UAE. Association for Computational Linguistics.
  20. 20.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022b. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  21. 21.John Hewitt and Percy Liang. 2019. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743.
  22. 22.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, Xian Li, Brian O’Horo, Gabriel Pereyra, Jeff Wang, Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, and Ves Stoyanov. 2023. Opt-iml: Scaling language model instruction meta learning through the lens of generalization.
  23. 23.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  24. 24.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  25. 25.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know.
  26. 26.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge. arXiv preprint arXiv:2211.08411.
  27. 27.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics, 7:453–466.
  28. 28.Rémi Leblond, Jean-Baptiste Alayrac, Laurent Sifre, Miruna Pislar, Lespiau Jean-Baptiste, Ioannis Antonoglou, Karen Simonyan, and Oriol Vinyals. 2021. Machine translation decoding beyond beam search. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8410–8434, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  29. 29.Belinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, and Jacob Andreas. 2022. Quantifying adaptability in pre-trained language models with 500 tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4696–4715, Seattle, United States. Association for Computational Linguistics.
  30. 30.Jinzhi Liao, Xiang Zhao, Jianming Zheng, Xinyi Li, Fei Cai, and Jiuyang Tang. 2022. Ptau: Prompt tuning for attributing unanswerable questions. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, page 1219–1229, New York, NY, USA. Association for Computing Machinery.
  31. 31.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint arXiv:2212.10511.
  32. 32.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.
  33. 33.Clara Meister, Ryan Cotterell, and Tim Vieira. 2020. If beam search is the answer, what was the question? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2173–2185, Online. Association for Computational Linguistics.
  34. 34.Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-editing memory in a transformer. In The International Conference on Learning Representations (ICLR).
  35. 35.Timothee Mickus, Denis Paperno, and Mathieu Constant. 2022. How to dissect a Muppet: The structure of transformer embedding spaces. Transactions of the Association for Computational Linguistics, 10:981–996.
  36. 36.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2022a. Reframing instructional prompts to GPTk’s language. In Findings of the Association for Computational Linguistics: ACL 2022, pages 589–612, Dublin, Ireland. Association for Computational Linguistics.
  37. 37.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022b. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  38. 38.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  39. 39.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  40. 40.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  41. 41.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256.
  42. 42.Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell. 2022a. Linear guardedness and its implications. arXiv preprint arXiv:2210.10012.
  43. 43.Shauli Ravfogel, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021. Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 194–209.
  44. 44.Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell. 2022b. Linear adversarial concept erasure. In International Conference on Machine Learning, pages 18400–18421. PMLR.
  45. 45.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  46. 46.Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. 2022. Confident adaptive language modeling. In Advances in Neural Information Processing Systems (NeurIPS).
  47. 47.Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. 2020. The right tool for the job: Matching model and instance complexities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6640–6651, Online. Association for Computational Linguistics.
  48. 48.Aviv Slobodkin, Avi Caciularu, Eran Hirsch, and Ido Dagan. 2023. Dont add, dont miss: Effective content preserving generation from pre-selected text spans.
  49. 49.Aviv Slobodkin, Leshem Choshen, and Omri Abend. 2021. Mediators in determining what processing BERT performs first. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 86–93, Online. Association for Computational Linguistics.
  50. 50.Elior Sulem, Jamaal Hay, and Dan Roth. 2021. Do we know what we don’t know? studying unanswerable questions beyond SQuAD 2.0. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4543–4548, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  51. 51.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multi-hop questions via single-hop question composition. Transactions of the Association for Computational Linguistics.
  52. 52.Elena Voita, Rico Sennrich, and Ivan Titov. 2019. The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4396–4406, Hong Kong, China. Association for Computational Linguistics.
  53. 53.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023a. Poisoning language models during instruction tuning. arXiv preprint arXiv:2305.00944.
  54. 54.David Wan, Mengwen Liu, Kathleen McKeown, Markus Dreyer, and Mohit Bansal. 2023b. Faithfulness-aware decoding strategies for abstractive summarization. arXiv preprint arXiv:2303.03278.
  55. 55.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  56. 56.Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2023. " according to..." prompting language models improves quoting from pre-training data. arXiv preprint arXiv:2305.13252.
  57. 57.Gian Wiher, Clara Meister, and Ryan Cotterell. 2022. On decoding strategies for neural text generators. Transactions of the Association for Computational Linguistics, 10:997–1012.
  58. 58.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don’t know?
  59. 59.Haichao Zhu, Li Dong, Furu Wei, Wenhui Wang, Bing Qin, and Ting Liu. 2019. Learning to ask unanswerable questions for machine reading comprehension. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4238–4248, Florence, Italy. Association for Computational Linguistics.

Citation

MLA
Slobodkin, A., et al. “The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 3607–25, https://doi.org/10.18653/v1/2023.emnlp-main.220.
APA
Slobodkin, A., Goldman, O., Caciularu, A., Dagan, I., & Ravfogel, S. (2023). The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3607–3625. https://doi.org/10.18653/v1/2023.emnlp-main.220
Chicago
Slobodkin, A., O. Goldman, A. Caciularu, I. Dagan, and S. Ravfogel. 2023. “The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3607–25. https://doi.org/10.18653/v1/2023.emnlp-main.220.
Harvard
Slobodkin, A. et al. (2023) “The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3607–3625. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.220.
Vancouver
1. Slobodkin A, Goldman O, Caciularu A, Dagan I, Ravfogel S (2023) The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3607–3625

BibTeX

@inproceedings{slobodkin-etal-2023-curious,
    title = "The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models",
    author = "Slobodkin, Aviv  and
      Goldman, Omer  and
      Caciularu, Avi  and
      Dagan, Ido  and
      Ravfogel, Shauli",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.220/",
    doi = "10.18653/v1/2023.emnlp-main.220",
    pages = "3607--3625"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/