Inside-Out: Hidden Factual Knowledge in LLMs
Jonathan HerzigEran OfekHadas OrgadZorik GekhmanIdan SzpektorRoi ReichartYonatan BelinkovEyal Ben-David
Demonstrates that large language models encode an average of 40% more factual knowledge internally than they express in their outputs, exposing a fundamental limitation in scaling test-time compute through repeated sampling because certain known facts are never generated.
Large language models (LLMs) are widely deployed for knowledge-intensive applications, yet fundamental questions remain about the nature and reliability of their factual recall. Traditional benchmarks often evaluate a model based on a single generated response, obscuring whether the model genuinely lacks information or simply failed to output it during standard generation. Understanding whether models store more factual knowledge in their internal parameters than they express in their visible outputs is critical for improving model performance, advancing interpretability, and mitigating safety risks tied to unexpressed or suddenly surfacing data.
The article establishes a formal computational framework to define and quantify factual knowledge in LLMs and evaluates whether models systematically harbor "hidden knowledge." Specifically, it demonstrates that models consistently encode more factual knowledge in their intermediate internal computations than they express through observable token probabilities and standard generation.
To conduct this evaluation, the researchers designed a closed-book question-answering benchmark using approximately 1,700 unambiguous, entity-centric questions across four factual relations (such as authors and spouses). For each question, they gathered candidate answers by generating an initial greedy response and sampling 1,000 additional responses, which were categorized for correctness using an automated LLM judge verified against human annotations. The study evaluated three open-weight models in the 7-to-9 billion parameter range (Llama-3-8B, Mistral-7B, and Gemma-2-9B) alongside a larger 32-billion parameter model (Qwen3-32B). Knowledge was quantified as the model's ability to rank correct answers above incorrect ones. The researchers compared external scoring methods—which rely on observable output probabilities and direct prompting—against an internal scoring method using a linear classifier (probe) trained on the models' intermediate hidden states.
The findings reveal that models consistently possess substantial hidden factual knowledge, with internal scoring outperforming all observable external methods across every tested model and relation by an average relative margin of 40%. The magnitude of this gap varied by architecture, reaching 57% in Gemma-2-9B and 14% in Llama-3-8B, and persisted at 12.5% in the larger 32-billion parameter model. Most notably, the article identified an extreme failure mode: in 7.2% of all questions, models possessed perfect internal knowledge of the correct answer and ranked it above all incorrect alternatives, yet failed to generate it even once across 1,000 repeated sampling attempts. Furthermore, leveraging internal representations to select the best answer among 1,000 generated candidates improved overall question-answering accuracy by an average of 12% over greedy decoding.
These results demonstrate a fundamental bottleneck in the standard autoregressive generation and decoding process of modern LLMs, which functions somewhat analogously to a human "tip-of-the-tongue" state. In practical terms, this constraint significantly limits the effectiveness of scaling test-time compute through repeated sampling alone. The findings indicate that an additional 40% relative performance gain remains inaccessible simply because current sampling mechanisms cannot surface correct candidate answers that the model internally knows.
Based on these findings, developers and organizations should not rely solely on generation likelihood or superficial sampling to assess or extract model knowledge. Instead, researchers and practitioners should invest in developing next-generation decoding algorithms that incorporate internal representations to surface hidden facts at inference time. Training methodologies should also explore loss functions and reinforcement learning reward designs that expose models to multiple valid answer formulations and prioritize factuality over stylistic output fluency.
These conclusions are supported by statistically significant results and controlled data splits designed to eliminate memorization artifacts. However, users should consider certain limitations: the high computational cost of sampling and evaluating thousands of answers limited the primary scope to 7-to-9 billion parameter models, and the evaluation focused on single-hop, entity-centric facts rather than complex multi-hop reasoning. While confidence in the existence of the internal-external knowledge gap is high, the precise magnitude may vary across different prompt designs, broader knowledge domains, and larger model scales.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). It establishes foundational techniques for discovering latent factual knowledge directly from internal activations independently of generated text outputs.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). It provides the seminal benchmark and methodology for probing whether pretrained language models store factual knowledge in their parameters.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). It highlights how surface-level prompt generation often fails to elicit the factual knowledge actually stored inside language models.
- Paper: Dissecting Recall of Factual Associations in Auto-Regressive Language Models, Mor Geva et al. (2023). It analyzes the internal layer-by-layer mechanistic pathways through which autoregressive transformers retrieve and assemble factual associations.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). It investigates how well language models can evaluate the validity of their own knowledge and calibrate their factual predictions.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). It builds on internal factual representations by using sparse autoencoder steering to resolve conflicts between internal memory and external context.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). It extends the study of internal model representations by using intermediate hidden states and circuits to predict generation failures.
