keyword
latent knowledge
Latent knowledge refers to the factual information, logical truths, and world representations that are implicitly encoded within the internal states and activations of a machine learning model, distinct from what the model explicitly generates in its surface-level outputs. In artificial intelligence systems such as large language models, a model may internally represent whether a statement is true, probable, or logically consistent even when its overt text generation produces errors, sycophantic responses, or hallucinations. Researchers identify and extract this implicit information through interpretability methods, such as probing activation spaces or applying directional interventions, allowing them to measure what a model internally represents, evaluate its underlying confidence, and encourage more truthful and aligned behavior without relying solely on supervised output training.
3 items

Distinguishing the Knowable from the Unknowable with Language Models
Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, Benjamin L. Edelman
Why you should read this
Demonstrates that language models internally separate reducible epistemic uncertainty from inherent aleatoric entropy, enabling simple probes and unsupervised methods to accurately detect when an uncertain prediction is caused by a lack of knowledge rather than inherent ambiguity.
We study the feasibility of identifying epistemic uncertainty (reflecting a lack of knowledge), as opposed to aleatoric uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over free-form text. In the absence of ground-truth probabilities, we explore a setting where, in order to (approximately) disentangle a given LLM’s uncertainty, a significantly larger model stands in as a proxy for the ground truth. We show that small linear probes trained on the embeddings of frozen, pretrained models accurately predict when larger models will be more confident at the token level and that probes trained on one text domain generalize to others. Going further, we propose a fully unsupervised method that achieves non-trivial accuracy on the same task. Taken together, we interpret these results as evidence that LLMs naturally contain internal representations of different types of uncertainty that could potentially be leveraged to devise more informative indicators of model confidence in diverse practical settings. Code can be found at: https://github.com/KempnerInstitute/llm_uncertainty
Added
2026-10-05

Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt
Why you should read this
Introduces an unsupervised technique for extracting truthful latent knowledge directly from language model activations by enforcing logical consistency, enabling accurate question answering even when models are prompted to generate false outputs.
Existing techniques for training language models can be misaligned with the truth: if we train models with imitation learning, they may reproduce errors that humans make; if we train them to generate text that humans rate highly, they may output errors that human evaluators can't detect. We propose circumventing this issue by directly finding latent knowledge inside the internal activations of a language model in a purely unsupervised way. Specifically, we introduce a method for accurately answering yes-no questions given only unlabeled model activations. It works by finding a direction in activation space that satisfies logical consistency properties, such as that a statement and its negation have opposite truth values. We show that despite using no supervision and no model outputs, our method can recover diverse knowledge represented in large language models: across 6 models and 10 question-answering datasets, it outperforms zero-shot accuracy by 4\% on average. We also find that it cuts prompt sensitivity in half and continues to maintain high accuracy even when models are prompted to generate incorrect answers. Our results provide an initial step toward discovering what language models know, distinct from what they say, even when we don't have access to explicit ground truth labels.
Added
2026-09-26

Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Kenneth Li, Oam Patel, Fernanda Viégas, H. Pfister, M. Wattenberg
Why you should read this
Introduces Inference-Time Intervention, a data-efficient and minimally invasive technique that doubles the truthfulness of large language models on the TruthfulQA benchmark by shifting internal activations along truth-correlated directions during generation.
We introduce Inference-Time Intervention (ITI), a technique designed to enhance the “truthfulness” of large language models (LLMs). ITI operates by shifting model activations during inference, following a set of directions across a limited number of attention heads. This intervention significantly improves the performance of LLaMA models on the TruthfulQA benchmark. On an instruction-finetuned LLaMA called Alpaca, ITI improves its truthfulness from 32.5% to 65.1%. We identify a trade-off between truthfulness and helpfulness and demonstrate how to balance it by tuning the intervention strength. ITI is minimally invasive and computationally inexpensive. Moreover, the technique is data efficient: while approaches like RLHF require extensive annotations, ITI locates truthful directions using only few hundred examples. Our findings suggest that LLMs may have an internal representation of the likelihood of something being true, even as they produce falsehoods on the surface. Code: https://github.com/likenneth/honest_llama.
Added
2026-09-25
