The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models
Aviv SlobodkinOmer GoldmanAvi CaciularuIdo DaganShauli Ravfogel
Reveals that large language models internally encode whether a question is answerable in their hidden states even while generating hallucinatory answers, showing that this linearly separable answerability signal can be extracted to curb overconfident errors.
Large language models frequently generate confident but incorrect responses when faced with unanswerable questions. In real-world applications such as search engines and automated assistants, these fabricated answers create substantial compliance, safety, and reliability risks. The article investigates whether language models inherently track the answerability of a question within their internal representations even while generating incorrect text, and whether this internal state can be used to prevent such errors.
The researchers evaluated three instruction-tuned models across three reading comprehension benchmarks containing both answerable and unanswerable questions. They applied four distinct analytical techniques: modifying prompts to explicitly give models permission to abstain, inspecting alternative potential answers during text generation, training simple linear classifiers on the models' internal hidden layers to predict answerability, and mathematically erasing the identified answerability features to verify their direct influence on model behavior.
The findings show that language models internally recognize when a question cannot be answered from the provided text. First, simply including instructions that explicitly permit abstention improved unanswerability detection performance by up to 80 points, with overall question-answering accuracy increasing by more than 50 points. Second, when standard models fabricated an answer, correct abstention responses were typically present as lower-probability candidates in the generated beam. Third, a basic linear classifier trained on the internal state of the first generated word achieved over 75% accuracy across all models and benchmarks, demonstrating that answerability is clearly organized in the model's internal data space. Finally, erasing this internal feature substantially degraded performance, confirming its functional importance.
These results indicate that overconfident inaccuracies are not caused by a failure of model comprehension, but rather by generation mechanisms that prioritize generating a definitive answer over admitting ignorance. Organizations deploying language models should immediately incorporate explicit abstention instructions into system prompts. Furthermore, engineering teams can implement lightweight internal classifiers or search candidate response sets to filter out unanswerable queries before responses reach end users.
The primary limitations include a focus on reading comprehension tasks within specific text contexts rather than open-domain knowledge queries, as well as testing a selected set of model architectures. Future work should assess how these internal mechanisms function in open-domain environments and evaluate larger model sizes.
- Paper: Know What You Don't Know: Unanswerable Questions for SQuAD, Pranav Rajpurkar et al. (2018). SQuAD 2.0 established the answerable-versus-unanswerable reading-comprehension setting that makes this paper’s answerability results interpretable.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its evidence that models can estimate whether their own answers are correct provides essential groundwork for this paper’s claim that answerability is represented internally.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). Its method for extracting latent knowledge from hidden activations helps explain the probing approach used to detect answerability in internal states.
- Paper: Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?, Kevin Liu et al. (2023). Its distinction between model outputs and probe-based truthfulness judgments prepares readers for this paper’s analysis of confident answers that conflict with internal answerability signals.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). It extends internal-state failure detection beyond answerability, using hidden representations and attention signals to predict when a model’s generated response is wrong.
- Paper: Unfamiliar Finetuning Examples Control How Language Models Hallucinate, Katie Kang et al. (2025). It carries the abstention insight into training, showing how unfamiliar examples labeled with “I don’t know” can teach models when to abstain.
- Paper: RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models, Aashiq Muhamed et al. (2025). It continues the selective-refusal problem by developing a dynamic evaluation framework for whether models abstain appropriately when grounded context is uncertain.
- Paper: Inducing Artificial Uncertainty in Language Models, Sophia Hager et al. (2026). It advances internal uncertainty probing by training detectors on artificially induced uncertainty and testing whether they transfer to difficult questions.
