Discovering Latent Knowledge in Language Models Without Supervision

Collin BurnsHaotian YeDan KleinJacob Steinhardt

article2023ICLR626 citations

Introduces an unsupervised technique for extracting truthful latent knowledge directly from language model activations by enforcing logical consistency, enabling accurate question answering even when models are prompted to generate false outputs.

Listen

Current artificial intelligence systems trained on human text or feedback often generate factual errors, mirror common human misconceptions, or provide misleading outputs when incentivized to do so. As models are deployed in increasingly complex domains where human evaluators cannot easily verify correctness, conventional supervised alignment methods risk breaking down. The article evaluates whether it is possible to extract accurate, latent knowledge directly from the internal activations of large language models without using any labeled data, supervision, or generated text.

The researchers developed a method called Contrast-Consistent Search (CCS). The approach evaluates binary (yes-no) question pairs by pairing statements with their negations, mapping model hidden states to probabilities, and enforcing basic logical consistency principles—specifically, that a statement and its negation cannot both be true, and one must be true. The method was evaluated across six language models (such as GPT-J, T5, and DeBERTa) and ten diverse classification and question-answering benchmarks covering sentiment, factual reasoning, and textual entailment.

The evaluation yielded several key findings. First, CCS achieved an average classification accuracy of 71.2% across benchmarks, outperforming standard calibrated zero-shot baselines by 4% on average without requiring ground-truth labels. Second, the method reduced sensitivity to prompt phrasing, cutting the standard deviation of accuracy across different prompt variations in half compared to zero-shot outputs. Third, CCS proved robust against deliberate manipulation: when models were primed with misleading context that caused zero-shot accuracy to drop by 9.5%, CCS maintained high accuracy. Finally, the learned truth representations transferred effectively across entirely different tasks and remained recoverable from intermediate layers even when raw text outputs were uninformative.

These findings indicate that language models internally encode structured, task-agnostic representations of truth that are distinct from, and often more accurate than, their generated text. This demonstrates that external human supervision may not be strictly necessary to identify what an AI model actually "knows." For safety and compliance oversight, discovering latent internal knowledge offers a potential mechanism to detect model deception or errors that bypass surface-level human monitoring.

Organizations developing or auditing high-stakes AI applications should explore internal representation probing alongside traditional output-level evaluations. However, further technical development is required before deploying these methods for production oversight. Researchers and practitioners must expand the technique beyond binary choices to open-ended and non-clear-cut statements, improve calibration, and formally test the method against active strategic deception. Confidence in the reported results is high, supported by statistically significant gains across multiple models and benchmarks, though applicability remains bounded to domains where models have formed distinct internal truth representations.

  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This paper establishes foundational empirical evidence on language model calibration and self-evaluating internal knowledge, which provides essential grounding for discovering latent knowledge without supervision.
  • Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Understanding how standard training leads language models to mimic human falsehoods motivates the need for unsupervised methods to recover truthful internal activations.
  • Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This work introduces the paradigm of probing pretrained language models for stored relational facts, establishing early benchmarks for measuring factual representations.
  • Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This study demonstrates how surface prompting can fail to elicit stored model knowledge, highlighting the core motivation for probing latent representations directly.
  • Paper: BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions, Christopher Clark et al. (2019). This paper presents the standard dataset and benchmark for evaluating yes/no question answering, serving as a core evaluation format for latent knowledge probes.
Cover for Discovering Latent Knowledge in Language Models Without Supervision

Abstract

Existing techniques for training language models can be misaligned with the truth: if we train models with imitation learning, they may reproduce errors that humans make; if we train them to generate text that humans rate highly, they may output errors that human evaluators can't detect. We propose circumventing this issue by directly finding latent knowledge inside the internal activations of a language model in a purely unsupervised way. Specifically, we introduce a method for accurately answering yes-no questions given only unlabeled model activations. It works by finding a direction in activation space that satisfies logical consistency properties, such as that a statement and its negation have opposite truth values. We show that despite using no supervision and no model outputs, our method can recover diverse knowledge represented in large language models: across 6 models and 10 question-answering datasets, it outperforms zero-shot accuracy by 4% on average. We also find that it cuts prompt sensitivity in half and continues to maintain high accuracy even when models are prompted to generate incorrect answers. Our results provide an initial step toward discovering what language models know, distinct from what they say, even when we don't have access to explicit ground truth labels.

Table of Contents

  • 1 Introduction
  • 2 Problem Statement and Framework
  • 2.1 Problem: Discovering Latent Knowledge
  • 2.2 Method: Contrast-Consistent Search
  • 3 Results
  • 3.1 Experimental Setup
  • 3.2 Evaluating CCS
  • 3.2.1 CCS Outperforms Zero-Shot
  • 3.2.2 CCS Is Robust To Misleading Prompts
  • 3.3 Analyzing CCS
  • 3.3.1 CCS Finds A Task-Agnostic Representation of Truth
  • 3.3.2 CCS Does Not Just Recover Model Outputs
  • 3.3.3 Truth is a Salient Feature
  • 4 Related Work
  • 5 Discussion
  • 5.1 Limitations and Future Work
  • 5.2 Conclusion
  • References
  • A Identifying Clusters
  • B Misleading Prefix Details
  • B.1 Misleading Prefix
  • B.1.1 Evaluating The Effect of The Misleading Prefix
  • C Masked Language Modeling Results
  • D Complete Sample Complexity Results
  • E Complete Transfer Results
  • F Complete Intermediate Representations Results
  • G CCS and CRC Implementation Details
  • G.1 Normalization
  • G.2 Contrast Pairs
  • G.3 CCS
  • G.4 CRC: Top Principal Component
  • G.5 CRC: Bimodal Salience Search
  • H Statistical Significance
  • I Dataset Setup
  • I.1 Tokenization
  • I.2 Datasets
  • I.2.1 AG_News
  • I.2.2 Amazon_polarity
  • I.2.3 BOOLQ
  • I.2.4 COPA
  • I.2.5 DBpedia_14
  • I.2.6 IMDB
  • I.2.7 PIQA
  • I.2.8 QNLI
  • I.2.9 RTE
  • I.2.10 Story_Cloze

Knowls

  1. Knowl 1 — Contrast-Consistent Search (CCS) Framework

    model/method

    Contrast-Consistent Search (CCS) is an unsupervised method designed to elicit latent factual and procedural knowledge from the internal activations of a pretrained language model without requiring ground-truth labels, fine-tuning model parameters, or relying on model text generation.

    Given a set of binary or yes-no questions q1,…,qnq_1, \dots, q_n, CCS operates through the following steps:

    1. Contrast Pair Construction: Each question qiq_i is formatted into two opposing natural language statements: xi+x_i^+ (representing the positive answer, such as "Yes") and xi−x_i^- (representing the negative answer, such as "No").
    2. Feature Extraction and Normalization: For each statement, an internal hidden representation is extracted from a specified layer of the language model, yielding ϕ(xi+),ϕ(xi−)∈Rd\phi(x_i^+), \phi(x_i^-) \in \mathbb{R}^d, typically at the final token position. To eliminate superficial differences caused by specific answer tokens (such as "Yes" vs. "No"), the representations are centered independently: ϕ~(xi+)=ϕ(xi+)−μ+,ϕ~(xi−)=ϕ(xi−)−μ−controls\tilde{\phi}(x_i^+) = \phi(x_i^+) - \mu^+, \quad \tilde{\phi}(x_i^-) = \phi(x_i^-) - \mu^- controls where μ+,μ−∈Rd\mu^+, \mu^- \in \mathbb{R}^d are the sample means of {ϕ(xi+)}i=1n\{\phi(x_i^+)\}_{i=1}^n and {ϕ(xi−)}i=1n\{\phi(x_i^-)\}_{i=1}^n, respectively. The features are also scaled by dividing by their average Euclidean norm multiplied by d\sqrt{d}.
    3. Probability Mapping: A probe parameterized by weight θ∈Rd\theta \in \mathbb{R}^d and bias b∈Rb \in \mathbb{R} maps normalized representations to truth probabilities via a logistic sigmoid: pθ,b(ϕ~)=σ(θTϕ~+b)p_{\theta, b}(\tilde{\phi}) = \sigma(\theta^T \tilde{\phi} + b)
    4. Unsupervised Optimization: The probe parameters θ,b\theta, b are optimized by gradient descent on an unsupervised objective enforcing logical consistency across negations and confidence.
    5. Inference: For a question qiq_i, the estimated probability that the positive statement xi+x_i^+ is true is computed as: p~(qi)=12(pθ,b(ϕ~(xi+))+(1−pθ,b(ϕ~(xi−))))\tilde{p}(q_i) = \frac{1}{2}\left(p_{\theta, b}(\tilde{\phi}(x_i^+)) + (1 - p_{\theta, b}(\tilde{\phi}(x_i^-)))\right) The model predicts the positive answer if p~(qi)>0.5\tilde{p}(q_i) > 0.5.
  2. Knowl 2 — Loss Function Formulation for Contrast-Consistent Search

    equation

    The Contrast-Consistent Search (CCS) objective finds an internal direction in representation space satisfying logical consistency constraints. For a probe pθ,b(ϕ~)=σ(θTϕ~+b)p_{\theta, b}(\tilde{\phi}) = \sigma(\theta^T \tilde{\phi} + b) applied to normalized contrast activations (ϕ~(xi+),ϕ~(xi−))(\tilde{\phi}(x_i^+), \tilde{\phi}(x_i^-)) corresponding to question qiq_i, the loss is:

    LCCS(θ,b)=1n∑i=1n[Lconsistency(θ,b;qi)+Lconfidence(θ,b;qi)]L_{\text{CCS}}(\theta, b) = \frac{1}{n} \sum_{i=1}^n \left[ L_{\text{consistency}}(\theta, b; q_i) + L_{\text{confidence}}(\theta, b; q_i) \right]

    where:

    1. Consistency Loss penalizes deviations from the law of negation (probabilities of contradictory statements must sum to 1): Lconsistency(θ,b;qi)=[pθ,b(ϕ~(xi+))−(1−pθ,b(ϕ~(xi−)))]2L_{\text{consistency}}(\theta, b; q_i) = \left[ p_{\theta, b}(\tilde{\phi}(x_i^+)) - (1 - p_{\theta, b}(\tilde{\phi}(x_i^-))) \right]^2

    2. Confidence Loss penalizes uninformative probabilities near 0.5 by enforcing the law of excluded middle (at least one statement in a contrast pair must have high probability of being true, so the smaller probability must be close to 0): Lconfidence(θ,b;qi)=min⁡{pθ,b(ϕ~(xi+)), pθ,b(ϕ~(xi−))}2L_{\text{confidence}}(\theta, b; q_i) = \min\left\{ p_{\theta, b}(\tilde{\phi}(x_i^+)),\, p_{\theta, b}(\tilde{\phi}(x_i^-)) \right\}^2

    Both loss components are required: LconsistencyL_{\text{consistency}} alone admits a degenerate trivial solution p(x+)=p(x−)=0.5p(x^+) = p(x^-) = 0.5, while LconfidenceL_{\text{confidence}} alone admits degenerate solutions such as p(x+)=p(x−)=0p(x^+) = p(x^-) = 0.

  3. Knowl 3 — Unsupervised Cluster Disambiguation via Logical Conjunctions and Disjunctions

    theoretical result

    Because the unsupervised objective LCCSL_{\text{CCS}} is symmetric with respect to truth and falsehood, it identifies a 1D subspace that clusters true and false statements into two separated groups but cannot inherently determine which cluster corresponds to truth and which corresponds to falsehood (sign ambiguity).

    If the underlying language model is capable of processing logical connectives, this sign ambiguity can be resolved in a fully unsupervised manner:

    • Let xx and x′x' be statements identified by CCS as belonging to opposite clusters (meaning one is true and the other is false).
    • The conjunction x∧x′x \land x' must be false.
    • The disjunction x∨x′x \lor x' must be true.

    Evaluating the probe's predictions on the compound statements x∧x′x \land x' and x∨x′x \lor x' unambiguously identifies which of the two clusters represents the truth direction without requiring any ground-truth labels.

  4. Knowl 4 — Contrastive Representation Clustering (CRC: TPC and BSS)

    model/method

    Contrastive Representation Clustering (CRC) comprises alternative unsupervised methods that identify truth representations by clustering difference vectors of normalized contrast pairs ci=ϕ~(xi+)−ϕ~(xi−)c_i = \tilde{\phi}(x_i^+) - \tilde{\phi}(x_i^-), based on the observation that the primary semantic variation between a statement and its negation is its truth value.

    Two variants are defined:

    1. Top Principal Component (TPC): Principal Component Analysis (PCA) is performed on the set of contrast vectors {ci}i=1n\{c_i\}_{i=1}^n. Predictions are made by projecting cic_i onto the first principal component and thresholding at 0: y^i=I(θTPCTci>0)\hat{y}_i = \mathbb{I}(\theta_{\text{TPC}}^T c_i > 0)
    2. Bimodal Salience Search (BSS): To address the issue that variance in TPC is sensitive to artificial feature scaling across layers, BSS minimizes an unsupervised variance ratio analogous to Linear Discriminant Analysis (LDA) that is invariant to overall scale: LBSS(θ)=var{θTci∣θTci<0}+var{θTci∣θTci≥0}var{θTci}L_{\text{BSS}}(\theta) = \frac{\text{var}\{\theta^T c_i \mid \theta^T c_i < 0\} + \text{var}\{\theta^T c_i \mid \theta^T c_i \ge 0\}}{\text{var}\{\theta^T c_i\}} Optimization is performed using projected gradient descent over the unit sphere ∥θ∥2=1\|\theta\|_2 = 1 with the Adam optimizer.
  5. Knowl 5 — Experimental Evaluation Suite for Discovering Latent Knowledge

    experimental setup

    The benchmark evaluates unsupervised knowledge extraction across six language model families and ten datasets converted into binary classification or yes-no question answering format.

    • Models:
      • Encoder-Decoder: T5-11B, UnifiedQA-11B, T0-11B.
      • Autoregressive: GPT-J (6B parameters).
      • Encoder-only: RoBERTa-large, DeBERTa-v2-xxlarge (both evaluated using MNLI-finetuned checkpoints for standardized zero-shot logit extraction).
    • Datasets (10 total):
      • Sentiment: IMDB, Amazon Polarity.
      • Topic Classification: AG-News, DBpedia-14.
      • Natural Language Inference (NLI): RTE, QNLI.
      • Causal/Story Reasoning: COPA, Story-Cloze.
      • Question Answering: BoolQ.
      • Commonsense Reasoning: PIQA.
    • Data Protocol: 1,000 balanced examples are subsampled per dataset (500 for COPA) and split 60% unsupervised training / 40% test. To assess prompt sensitivity, 8 to 13 distinct prompt templates (averaging 9 per dataset) are used, totaling approximately 180,000 evaluated instances.
    • CCS Training: Optimized with AdamW (learning rate η=0.01\eta = 0.01, E=1000E = 1000 epochs) across 10 random restarts, selecting the run with the lowest unsupervised loss LCCSL_{\text{CCS}}.
  6. Knowl 6 — Zero-Shot Accuracy and Prompt Sensitivity Comparison

    data/table

    Across five non-contaminated models (excluding T0*, which was pretrained on 9 of the 10 evaluation datasets), CCS achieves an average test accuracy of 71.2%, outperforming calibrated zero-shot (67.2%) by 4.0% and uncalibrated zero-shot (62.8%) by 8.4%. Furthermore, CCS reduces prompt sensitivity, cutting the average standard deviation across prompts from 6.1% to 3.2%.

    Method RoBERTa DeBERTa GPT-J T5 UQA T0* Mean*
    0-shot 60.1(5.7) 68.6(8.2) 53.2(5.2) 55.4(5.7) 76.8(9.6) 87.9(4.8) 62.8(6.9)
    Calibrated 0-shot 64.3(6.2) 76.3(6.0) 56.0(5.2) 58.8(6.1) 80.4(7.1) 90.5(2.7) 67.2(6.1)
    CCS 62.1(4.1) 78.5(3.8) 61.7(2.5) 71.5(3.0) 82.1(2.7) 77.6(3.3) 71.2(3.2)
    CCS (All Data) 60.1(3.7) 77.1(4.1) 62.1(2.3) 72.7(6.0) 84.8(2.6) 84.8(3.7) 71.5(3.7)
    LR (Ceiling) 79.8(2.5) 86.1(2.2) 78.1(2.3) 84.6(3.1) 89.8(1.9) 90.5(2.1) 83.7(2.4)

    Values represent accuracy percentages averaged over all prompts and datasets, with standard deviations across prompts in parentheses. Mean excludes T0* due to dataset overlap in T0 pretraining. Calibrated zero-shot adjusts classification logits with a shift parameter γ\gamma to balance predictions 50/50. LR (Ceiling) is a supervised logistic regression probe trained on labeled activation pairs (ϕ~(x+),ϕ~(x−))(\tilde{\phi}(x^+), \tilde{\phi}(x^-)).

  7. Knowl 7 — Robustness of CCS to Misleading Few-Shot Contexts

    empirical result

    When language models are prompted with a misleading few-shot prefix containing deliberately incorrect or nonsensical question-answer demonstrations, model generation behavior imitates the context errors, causing significant drops in zero-shot accuracy. However, Contrast-Consistent Search (CCS) probe accuracy remains robust:

    • On UnifiedQA, prepending the misleading prefix causes calibrated zero-shot accuracy to fall from 80.4% to 70.9% (a 9.5% absolute drop).
    • Under the identical misleading conditions, CCS accuracy on UnifiedQA increases slightly from 82.1% to 83.8%.
    • Across all five non-contaminated models, calibrated zero-shot accuracy drops from 67.2% to 65.4% when misleading prefixes are added, whereas CCS retains an average accuracy of 71.2% (identical to the unperturbed baseline).

    This indicates that internal activations continue to linearly represent the true factual status of inputs even when output generation probabilities are corrupted by misleading context.

  8. Knowl 8 — Cross-Task Transfer of Contrast-Consistent Search Probes

    empirical result

    Probes trained via Contrast-Consistent Search on one dataset generalize effectively when evaluated on completely different tasks with different label spaces and domains (e.g., transferring from sentiment classification on Amazon to reading comprehension on BoolQ or NLI on RTE) without retraining.

    • Probes trained on simple tasks such as Amazon sentiment polarity achieve an average cross-task transfer accuracy of 71.8% across all other nine datasets, slightly exceeding the within-task non-transfer average accuracy of 71.2%.
    • Cross-task transfer accuracy is consistent across diverse source datasets (e.g., training on IMDB yields transfer performance comparable to training on DBpedia).

    These transfer dynamics demonstrate that the linear direction identified by CCS corresponds to a general, task-agnostic geometric representation of truth within the language model's latent activation space.

  9. Knowl 9 — Extraction of Truth from Intermediate Layers and Masked Language Models

    empirical result

    CCS recovers latent representations of truth that are distinct from, and decoupled from, final-layer output generation probabilities:

    1. Intermediate Layer Superiority: In encoder-decoder models like T5 and UnifiedQA, CCS probes trained on hidden states from intermediate (middle) encoder layers outperform those trained on final decoder layers, especially under misleading prompt prefixes where decoder accuracy drops from 81.0% to 73.5% while middle encoder layers remain robust.
    2. Masked Language Modeling Without Mask Tokens: When evaluated on DeBERTa-v2 (pretrained solely via Masked Language Modeling without NLI fine-tuning), without using any [MASK] tokens and with answer labels placed in the middle of the prompt (making final-token logits uninformative for next-token prediction), CCS achieves 93.7% accuracy on Amazon sentiment, compared to 71.6% for zero-shot output logit extraction.

    These results verify that CCS extracts latent knowledge from model representations rather than merely reading out surface-level generation logits.

  10. Knowl 10 — Sample Efficiency and Low-Dimensional Salience of Latent Truth

    empirical result

    The linear direction representing truth in language model activation space is highly salient and sample-efficient to discover:

    • CCS achieves near-asymptotic performance with very small numbers of unlabeled contrast pairs kk. On UnifiedQA, DeBERTa, and T0, training CCS on k=8k = 8 to k=64k = 64 unlabeled samples from a single prompt yields accuracy within 1–2% of training on the full dataset of 600 examples across all prompts.
    • Unsupervised PCA on contrast differences ci=ϕ~(xi+)−ϕ~(xi−)c_i = \tilde{\phi}(x_i^+) - \tilde{\phi}(x_i^-) (Top Principal Component / TPC) achieves a mean accuracy of 70.1% across models, matching zero-shot performance without any parameter optimization.
    • Bimodal Salience Search (BSS) achieves a mean accuracy of 70.2% across models.

    This confirms that truth representations naturally align with principal directions of high variance and bimodal separation in contrastive activation space.

  11. Knowl 11 — Limitations and Scope Conditions of Contrast-Consistent Search

    limitation

    Contrast-Consistent Search has several fundamental limitations and scope conditions:

    1. Linear Separability Assumption: CCS requires that the language model has computed the truth value of the input and that this truth value is linearly separable in activation space. If a model lacks the capacity to evaluate a statement or does not compute its truth during the forward pass, CCS cannot recover it.
    2. Binary and Clear-Cut Truth Constraint: The formulation strictly relies on binary contrast pairs (xi+,xi−)(x_i^+, x_i^-) with mutually exclusive, well-defined truth values; it does not directly apply to open-ended generation, continuous truth values, or ambiguous/subjective statements.
    3. Deception and Instrumental Lying: CCS was evaluated on standard datasets and misleading few-shot contexts, but was not tested on reinforcement learning agents actively trained to execute strategic deception or instrumental lies.
    4. Supervised Performance Gap: Across all tested models, a significant accuracy gap remains between unsupervised CCS (mean 71.2%) and the supervised Logistic Regression ceiling (mean 83.7%).

Coverage note — Specific prompt string templates for all 10 individual datasets (derived from PromptSource) and tokenization delimiters were omitted as implementation details in favor of the core methods, equations, tables, and empirical findings.

References

  1. 1.Amanda Askell, Yushi Bai, Anna Chen, Dawn Drain, Deep Ganguli, T. J. Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, John Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, and Jared Kaplan. A general language assistant as a laboratory for alignment. ArXiv, abs/2112.00861, 2021.
  2. 2.Yushi Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, T. J. Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv, abs/2204.05862, 2022.
  3. 3.Iz Beltagy, Arman Cohan, Robert Logan IV, Sewon Min, and Sameer Singh. Zero- and few-shot nlp with pretrained language models. In ACL, 2022.
  4. 4.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021.
  5. 5.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 7432–7439, 2020.
  6. 6.Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In NIPS, 2016.
  7. 7.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen A. Creel, Jared Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren E. Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas F. Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, O. Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir P. Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Benjamin Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, J. F. Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Robert Reich, Hongyu Ren, Frieda Rong, Yusuf H. Roohani, Camilo Ruiz, Jack Ryan, Christopher R’e, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishna Parasuram Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramer, Rose E. Wang, William Wang, ` Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei A. Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. On the opportunities and risks of foundation models. ArXiv, abs/2108.07258, 2021.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. ArXiv, abs/2005.14165, 2020.
  9. 9.Paul Christiano, Ajeya Cotra, and Mark Xu. Eliciting latent knowledge. Technical report, ARC, 2022. URL https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/edit.
  10. 10.Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. ArXiv, abs/1706.03741, 2017.
  11. 11.Paul Francis Christiano, Buck Shlegeris, and Dario Amodei. Supervising strong learners by amplifying weak experts. ArXiv, abs/1810.08575, 2018.
  12. 12.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  13. 13.Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders. Truthful ai: Developing and governing ai that does not lie. ArXiv, abs/2110.06674, 2021.
  14. 14.Meta FAIR, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra. Human-level play in the game of diplomacy by combining language models with strategic reasoning. Science, pp. eade9097, 2022.
  15. 15.Rory A. Fisher. The use of multiple measurements in taxonomic problems. Annals of Human Genetics, 7:179–188, 1936.
  16. 16.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. ArXiv, abs/2006.03654, 2021.
  17. 17.Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. Unsolved problems in ml safety. ArXiv, abs/2109.13916, 2021.
  18. 18.Geoffrey Irving, Paul Francis Christiano, and Dario Amodei. Ai safety via debate. ArXiv, abs/1805.00899, 2018.
  19. 19.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. Maieutic prompting: Logically consistent reasoning with recursive explanations. ArXiv, abs/2205.11822, 2022.
  20. 20.Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving. Alignment of language agents. ArXiv, abs/2103.14659, 2021.
  21. 21.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. Unifiedqa: Crossing format boundaries with a single qa system. In FINDINGS, 2020.
  22. 22.Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang goo Lee, Kang Min Yoo, and Taeuk Kim. Ground-truth labels matter: A deeper look into input-label demonstrations. ArXiv, abs/2205.12685, 2022.
  23. 23.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  24. 24.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Soren Auer, et al. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web, 6(2):167–195, 2015.
  25. 25.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: a research direction. ArXiv, abs/1811.07871, 2018.
  26. 26.Stephanie C. Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In ACL, 2022.
  27. 27.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys (CSUR), 2022.
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692, 2019.
  29. 29.Ilya Loshchilov and Frank Hutter. Fixing weight decay regularization in adam. ArXiv, abs/1711.05101, 2017.
  30. 30.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In ACL, 2022.
  31. 31.Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pp. 142–150, 2011.
  32. 32.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald. On faithfulness and factuality in abstractive summarization. ArXiv, abs/2005.00661, 2020.
  33. 33.Julian McAuley and Jure Leskovec. Hidden factors and hidden topics: understanding rating dimensions with review text. In Proceedings of the 7th ACM conference on Recommender systems, pp. 165–172, 2013.
  34. 34.Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nathan McAleese. Teaching language models to support answers with verified quotes. ArXiv, abs/2203.11147, 2022.
  35. 35.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. Metaicl: Learning to learn in context. ArXiv, abs/2110.15943, 2022a.
  36. 36.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? ArXiv, abs/2202.12837, 2022b.
  37. 37.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pp. 46–51, 2017.
  38. 38.Reiichiro Nakano, Jacob Hilton, S. Arun Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. Webgpt: Browser-assisted question-answering with human feedback. ArXiv, abs/2112.09332, 2021.
  39. 39.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  40. 40.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nathan McAleese, and Geoffrey Irving. Red teaming language models with language models. ArXiv, abs/2202.03286, 2022.
  41. 41.Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. ArXiv, abs/1910.10683, 2020.
  42. 42.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  43. 43.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series, 2011.
  44. 44.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric Michael Smith, Y.-Lan Boureau, and Jason Weston. Recipes for building an open-domain chatbot. In EACL, 2021.
  45. 45.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang A. Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M SAIFUL BARI, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Rose Biderman, Leo Gao, T. G. Owe Bers, Thomas Wolf, and Alexander M. Rush. Multitask prompted training enables zero-shot task generalization. ArXiv, abs/2110.08207, 2021.
  46. 46.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan J. Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback. ArXiv, abs/2009.01325, 2020.
  47. 47.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification. ArXiv, abs/1803.05355, 2018.
  48. 48.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  49. 49.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  50. 50.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. ArXiv, abs/2109.01652, 2022a.
  51. 51.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903, 2022b.
  52. 52.Laura Weidinger, John F. J. Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zachary Kenton, Sande Minnich Brown, William T. Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel. Ethical and social risks of harm from language models. ArXiv, abs/2112.04359, 2021.
  53. 53.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, and Jamie Brew. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771, 2019.
  54. 54.Wenpeng Yin, Nazneen Rajani, Dragomir Radev, Richard Socher, and Caiming Xiong. Universal natural language processing with limited annotations: Try few-shot textual entailment as a start. In EMNLP, 2020.
  55. 55.Xiang Zhang, Junbo Zhao, and Yann LeCun. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28:649–657, 2015.
  56. 56.Tony Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. ArXiv, abs/2102.09690, 2021.
  57. 57.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. In EMNLP, 2021.
  58. 58.Chunting Zhou, Junxian He, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. Prompt consistency for zero-shot task generalization. ArXiv, abs/2205.00049, 2022.

Citation

MLA
Burns, C., et al. “Discovering Latent Knowledge in Language Models Without Supervision”. arXiv, 2022, http://arxiv.org/abs/2212.03827v2.
APA
Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2022). Discovering Latent Knowledge in Language Models Without Supervision. arXiv. http://arxiv.org/abs/2212.03827v2
Chicago
Burns, C., H. Ye, D. Klein, and J. Steinhardt. 2022. “Discovering Latent Knowledge in Language Models Without Supervision”. arXiv. http://arxiv.org/abs/2212.03827v2.
Harvard
Burns, C. et al. (2022) “Discovering Latent Knowledge in Language Models Without Supervision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.03827v2.
Vancouver
1. Burns C, Ye H, Klein D, Steinhardt J (2022) Discovering Latent Knowledge in Language Models Without Supervision. arXiv

BibTeX

@article{burns2022discovering,
  title = {Discovering Latent Knowledge in Language Models Without Supervision},
  author = {Burns, Collin and Ye, Haotian and Klein, Dan and Steinhardt, Jacob},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.03827v2},
  eprint = {2212.03827}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors