Discovering Latent Knowledge in Language Models Without Supervision
Collin BurnsHaotian YeDan KleinJacob Steinhardt
Introduces an unsupervised technique for extracting truthful latent knowledge directly from language model activations by enforcing logical consistency, enabling accurate question answering even when models are prompted to generate false outputs.
Current artificial intelligence systems trained on human text or feedback often generate factual errors, mirror common human misconceptions, or provide misleading outputs when incentivized to do so. As models are deployed in increasingly complex domains where human evaluators cannot easily verify correctness, conventional supervised alignment methods risk breaking down. The article evaluates whether it is possible to extract accurate, latent knowledge directly from the internal activations of large language models without using any labeled data, supervision, or generated text.
The researchers developed a method called Contrast-Consistent Search (CCS). The approach evaluates binary (yes-no) question pairs by pairing statements with their negations, mapping model hidden states to probabilities, and enforcing basic logical consistency principles—specifically, that a statement and its negation cannot both be true, and one must be true. The method was evaluated across six language models (such as GPT-J, T5, and DeBERTa) and ten diverse classification and question-answering benchmarks covering sentiment, factual reasoning, and textual entailment.
The evaluation yielded several key findings. First, CCS achieved an average classification accuracy of 71.2% across benchmarks, outperforming standard calibrated zero-shot baselines by 4% on average without requiring ground-truth labels. Second, the method reduced sensitivity to prompt phrasing, cutting the standard deviation of accuracy across different prompt variations in half compared to zero-shot outputs. Third, CCS proved robust against deliberate manipulation: when models were primed with misleading context that caused zero-shot accuracy to drop by 9.5%, CCS maintained high accuracy. Finally, the learned truth representations transferred effectively across entirely different tasks and remained recoverable from intermediate layers even when raw text outputs were uninformative.
These findings indicate that language models internally encode structured, task-agnostic representations of truth that are distinct from, and often more accurate than, their generated text. This demonstrates that external human supervision may not be strictly necessary to identify what an AI model actually "knows." For safety and compliance oversight, discovering latent internal knowledge offers a potential mechanism to detect model deception or errors that bypass surface-level human monitoring.
Organizations developing or auditing high-stakes AI applications should explore internal representation probing alongside traditional output-level evaluations. However, further technical development is required before deploying these methods for production oversight. Researchers and practitioners must expand the technique beyond binary choices to open-ended and non-clear-cut statements, improve calibration, and formally test the method against active strategic deception. Confidence in the reported results is high, supported by statistically significant gains across multiple models and benchmarks, though applicability remains bounded to domains where models have formed distinct internal truth representations.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This paper establishes foundational empirical evidence on language model calibration and self-evaluating internal knowledge, which provides essential grounding for discovering latent knowledge without supervision.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Understanding how standard training leads language models to mimic human falsehoods motivates the need for unsupervised methods to recover truthful internal activations.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This work introduces the paradigm of probing pretrained language models for stored relational facts, establishing early benchmarks for measuring factual representations.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This study demonstrates how surface prompting can fail to elicit stored model knowledge, highlighting the core motivation for probing latent representations directly.
- Paper: BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions, Christopher Clark et al. (2019). This paper presents the standard dataset and benchmark for evaluating yes/no question answering, serving as a core evaluation format for latent knowledge probes.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). This work directly builds upon internal activation probing of truthfulness to develop an active intervention technique that steers generative model outputs toward truthful answers.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). This paper extends the study of internal representations by introducing lightweight circuits on hidden activations to predict model errors and hallucinations across reasoning tasks.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). This paper investigates the downstream risks of output misalignment identified in latent knowledge research, analyzing how models systematically mimic user biases rather than internal truth.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This study analyzes the disparity between internal decision factors and verbalized reasoning traces, further exploring the divergence between what models say and what they compute.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). This research explores conversational verification to detect factual errors without ground truth labels, offering an alternative black-box approach to unsupervised truth discovery.
