Inside-Out: Hidden Factual Knowledge in LLMs

Jonathan HerzigEran OfekHadas OrgadZorik GekhmanIdan SzpektorRoi ReichartYonatan BelinkovEyal Ben-David

article2025arXiv0 citationsSAC best paper award for Model Analysis and Interpretability

Demonstrates that large language models encode an average of 40% more factual knowledge internally than they express in their outputs, exposing a fundamental limitation in scaling test-time compute through repeated sampling because certain known facts are never generated.

Listen

Large language models (LLMs) are widely deployed for knowledge-intensive applications, yet fundamental questions remain about the nature and reliability of their factual recall. Traditional benchmarks often evaluate a model based on a single generated response, obscuring whether the model genuinely lacks information or simply failed to output it during standard generation. Understanding whether models store more factual knowledge in their internal parameters than they express in their visible outputs is critical for improving model performance, advancing interpretability, and mitigating safety risks tied to unexpressed or suddenly surfacing data.

The article establishes a formal computational framework to define and quantify factual knowledge in LLMs and evaluates whether models systematically harbor "hidden knowledge." Specifically, it demonstrates that models consistently encode more factual knowledge in their intermediate internal computations than they express through observable token probabilities and standard generation.

To conduct this evaluation, the researchers designed a closed-book question-answering benchmark using approximately 1,700 unambiguous, entity-centric questions across four factual relations (such as authors and spouses). For each question, they gathered candidate answers by generating an initial greedy response and sampling 1,000 additional responses, which were categorized for correctness using an automated LLM judge verified against human annotations. The study evaluated three open-weight models in the 7-to-9 billion parameter range (Llama-3-8B, Mistral-7B, and Gemma-2-9B) alongside a larger 32-billion parameter model (Qwen3-32B). Knowledge was quantified as the model's ability to rank correct answers above incorrect ones. The researchers compared external scoring methods—which rely on observable output probabilities and direct prompting—against an internal scoring method using a linear classifier (probe) trained on the models' intermediate hidden states.

The findings reveal that models consistently possess substantial hidden factual knowledge, with internal scoring outperforming all observable external methods across every tested model and relation by an average relative margin of 40%. The magnitude of this gap varied by architecture, reaching 57% in Gemma-2-9B and 14% in Llama-3-8B, and persisted at 12.5% in the larger 32-billion parameter model. Most notably, the article identified an extreme failure mode: in 7.2% of all questions, models possessed perfect internal knowledge of the correct answer and ranked it above all incorrect alternatives, yet failed to generate it even once across 1,000 repeated sampling attempts. Furthermore, leveraging internal representations to select the best answer among 1,000 generated candidates improved overall question-answering accuracy by an average of 12% over greedy decoding.

These results demonstrate a fundamental bottleneck in the standard autoregressive generation and decoding process of modern LLMs, which functions somewhat analogously to a human "tip-of-the-tongue" state. In practical terms, this constraint significantly limits the effectiveness of scaling test-time compute through repeated sampling alone. The findings indicate that an additional 40% relative performance gain remains inaccessible simply because current sampling mechanisms cannot surface correct candidate answers that the model internally knows.

Based on these findings, developers and organizations should not rely solely on generation likelihood or superficial sampling to assess or extract model knowledge. Instead, researchers and practitioners should invest in developing next-generation decoding algorithms that incorporate internal representations to surface hidden facts at inference time. Training methodologies should also explore loss functions and reinforcement learning reward designs that expose models to multiple valid answer formulations and prioritize factuality over stylistic output fluency.

These conclusions are supported by statistically significant results and controlled data splits designed to eliminate memorization artifacts. However, users should consider certain limitations: the high computational cost of sampling and evaluating thousands of answers limited the primary scope to 7-to-9 billion parameter models, and the evaluation focused on single-hop, entity-centric facts rather than complex multi-hop reasoning. While confidence in the existence of the internal-external knowledge gap is high, the precise magnitude may vary across different prompt designs, broader knowledge domains, and larger model scales.

Cover for Inside-Out: Hidden Factual Knowledge in LLMs

Abstract

This work presents a framework for assessing whether large language models (LLMs) encode more factual knowledge in their parameters than what they express in their outputs. While a few studies hint at this possibility, none has clearly defined or demonstrated this phenomenon. We first propose a formal definition of knowledge, quantifying it for a given question as the fraction of correct-incorrect answer pairs where the correct one is ranked higher. This gives rise to external and internal knowledge, depending on the information used to score individual answer candidates: either the model's observable token-level probabilities or its intermediate computations. Hidden knowledge arises when internal knowledge exceeds external knowledge. We then present a case study, applying this framework to three popular open-weights LLMs in a closed-book QA setup. Our results indicate that: (1) LLMs consistently encode more factual knowledge internally than what they express externally, with an average relative gap of 40%. (2) Surprisingly, some knowledge is so deeply hidden that a model can internally know an answer perfectly, yet fail to generate it even once, despite large-scale repeated sampling of 1,000 answers. This reveals fundamental limitations in the generation capabilities of LLMs, which (3) put a practical constraint on scaling test-time compute via repeated answer sampling in closed-book QA: significant performance improvements remain inaccessible because some answers are practically never sampled, yet if they were, we would be guaranteed to rank them first.

Table of Contents

  • 1 Introduction
  • 2 Hidden Knowledge
  • 2.1 Defining Knowledge Relative to an Answer Scoring Method
  • 2.2 Evidence of Hidden Knowledge
  • 2.3 Estimation
  • 3 Study Design
  • 3.1 Collecting the Set of Factual Triplets 𝒟={(𝐬i,𝐫i,𝐨i)}i=1n\mathcal{D}=\{(\mathbf{s}_{i},\mathbf{r}_{i},\mathbf{o}_{i})\}_{i=1}^{n}
  • 3.2 Approximating the Quantities of Interest
  • 4 Results
  • 4.1 Evidence of Hidden Knowledge in LLMs
  • 4.2 LLMs Can Fail to Generate Facts They Fully Know, Even After 1,000 Attempts
  • 4.3 A Case Study
  • 4.4 Increasing Test-Time Compute via Repeated Answer Sampling and Ranking in Closed-Book QA
  • 4.5 Hidden Knowledge in Larger Models
  • 5 Related Work
  • 6 Conclusion and Future Work
  • 7 Limitations
  • 8 Acknowledgements
  • References
  • A Appendix
  • A.1 Data Creation Process
  • A.2 QA Prompts
  • A.3 LLM Judge
  • A.4 Knowledge-aware Probe (on MM’s hidden states)
  • A.5 Training the Probe (on MM’s hidden states)
  • A.6 Evaluating Memorization in the Probe
  • A.7 Statistical Significance
  • A.8 External Scoring
  • A.8.1 𝐏⁡(𝐚|𝐪)\mathbf{P(a|q)} and 𝐏𝐧𝐨𝐫𝐦​(𝐚|𝐪)\mathbf{P_{norm}(a|q)}
  • A.8.2 P(True)
  • A.9 Extended Definition of Knowledge
  • A.10 Choosing the LLMs For Our Study
  • A.11 Analysis of 𝐊\mathbf{K} Values When Manually Adding The Gold Answer to 𝐀~​(𝐨)\mathbf{\tilde{A}(o)}
  • A.12 How 𝐊\mathbf{K} Affects Our Chances of Success in Inference Scaling?
  • A.13 Alternatives to QA Format

Knowls

  1. Knowl 1 — Pairwise Ranking Definition of Knowledge for Language Models

    definition

    For an autoregressive language model MM and a factual knowledge item represented as a subject-relation-object triplet (s,r,o)(s, r, o) (e.g., (France,capital,Paris)(\text{France}, \text{capital}, \text{Paris})), knowledge is defined relative to an answer scoring function SM:Q(s,r)×A~(o)→RS_M: Q(s,r) \times \tilde{A}(o) \to \mathbb{R}, which assigns a real-valued score to a question-answer pair (q,a)(q, a) using information derived from MM.

    Let the following sets be defined:

    • Q(s,r)Q(s, r): the set of paraphrased questions asking for entity oo given subject ss and relation rr.
    • A~(o)\tilde{A}(o): the set of all plausible answers having the same semantic entity type as oo.
    • A(o)⊆A~(o)A(o) \subseteq \tilde{A}(o): the set of all valid paraphrases and surface forms of the correct entity oo.
    • Ω(s,r,o):=A(o)×(A~(o)∖A(o))\Omega(s, r, o) := A(o) \times (\tilde{A}(o) \setminus A(o)): the set of all ordered pairs (a,a~)(a, \tilde{a}) containing a correct answer a∈A(o)a \in A(o) and an incorrect plausible answer a~∈A~(o)∖A(o)\tilde{a} \in \tilde{A}(o) \setminus A(o).

    The per-question knowledge score Kq(s,r,o;SM)K_q(s, r, o; S_M) quantifies the fraction of pairs in which the correct answer is scored strictly higher than the incorrect answer:

    Kq(s,r,o;SM)=1∣Ω(s,r,o)∣∑(a,a~)∈Ω(s,r,o)I(SM(q,a)>SM(q,a~))K_q(s, r, o; S_M) = \frac{1}{|\Omega(s, r, o)|} \sum_{(a, \tilde{a}) \in \Omega(s, r, o)} \mathbb{I}(S_M(q, a) > S_M(q, \tilde{a}))

    where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    The overall knowledge degree K(s,r,o;SM)K(s, r, o; S_M) across all question paraphrases is:

    K(s,r,o;SM)=1∣Q(s,r)∣∑q∈Q(s,r)Kq(s,r,o;SM)K(s, r, o; S_M) = \frac{1}{|Q(s, r)|} \sum_{q \in Q(s, r)} K_q(s, r, o; S_M)

    The binary indicator of full (perfect) knowledge K∗(s,r,o;SM)K^*(s, r, o; S_M) is:

    K∗(s,r,o;SM)=I(K(s,r,o;SM)=1)K^*(s, r, o; S_M) = \mathbb{I}(K(s, r, o; S_M) = 1)

  2. Knowl 2 — Criterion for Evidence of Hidden Knowledge in Language Models

    definition

    Let MM be a language model, SME\mathcal{S}_M^E the set of all external scoring functions that score an answer candidate (q,a)(q, a) relying exclusively on observable outputs from MM (such as token generation probabilities PM(a∣q)P_M(a \mid q), length-normalized probabilities Pnorm(a∣q)P_{\text{norm}}(a \mid q), or prompted verification token likelihoods PM("True"∣q,a)P_M(\text{"True"} \mid q, a)), and TMT_M an internal scoring function that leverages intermediate internal representations (such as hidden states hM(q,a)h_M(q, a) via a probing classifier).

    For a dataset D={(si,ri,oi)}i=1n\mathcal{D} = \{(s_i, r_i, o_i)\}_{i=1}^n of unique factual triplets, TMT_M provides empirical evidence that MM possesses hidden knowledge of D\mathcal{D} if the average knowledge score under TMT_M strictly exceeds the maximum average knowledge score under any external scoring function SM∈SMES_M \in \mathcal{S}_M^E by a statistically significant margin Δ>0\Delta > 0:

    1n∑i=1nK(si,ri,oi;TM)>max⁡SM∈SME(1n∑i=1nK(si,ri,oi;SM))+Δ\frac{1}{n} \sum_{i=1}^n K(s_i, r_i, o_i; T_M) > \max_{S_M \in \mathcal{S}_M^E} \left( \frac{1}{n} \sum_{i=1}^n K(s_i, r_i, o_i; S_M) \right) + \Delta

    where K(si,ri,oi;⋅)K(s_i, r_i, o_i; \cdot) is the pairwise ranking knowledge degree evaluated on triplet (si,ri,oi)(s_i, r_i, o_i).

  3. Knowl 3 — Empirical Evidence of Hidden Factual Knowledge Across Large Language Models

    empirical result

    Across 1,700 closed-book factual QA questions drawn from EntityQuestions across four distinct Wikidata relations—P26 (spouse), P176 (manufacturer), P264 (record label), and P50 (author)—internal scoring via a linear probe on model hidden states consistently and statistically significantly (p<0.05p < 0.05 under a paired tt-test) achieves higher knowledge scores KK and perfect knowledge scores K∗K^* than all external scoring methods (P(a∣q)P(a \mid q), Pnorm(a∣q)P_{\text{norm}}(a \mid q), and P(True)P(\text{True})).

    Key quantitative findings across evaluated models:

    • Across all 12 evaluated model-relation configurations, the average relative gap between the internal probe score and the strongest external scoring method is 40%40\%.
    • Gemma-2-9B-Instruct exhibits the largest average relative gap in KK (57%57\% relative increase over the best external function across relations, with relative improvements in K∗K^* reaching +1,500%+1,500\% on P264).
    • Mistral-7B-Instruct exhibits a 48%48\% average relative gap in KK, with relative gains in K∗K^* ranging from +121%+121\% to +250%+250\%.
    • Llama-3-8B-Instruct exhibits an average relative gap of 14%14\% in KK (gains in K∗K^* ranging from +8.5%+8.5\% to +18%+18\%), with P(True)P(\text{True}) closing part of the gap relative to probability-of-generation baselines.
    • In a 32B model, Qwen3-32B, the internal probe maintains a statistically significant relative gap of 12.5%12.5\% in KK and up to +47%+47\% in K∗K^* over the best external method, showing that hidden knowledge persists across model scale.
  4. Knowl 4 — Severe Generation Bottlenecks for Perfectly Encoded Internal Facts

    empirical result

    Language models frequently fail to generate factual answers that they internally represent with perfect accuracy, even after exhaustive repeated sampling:

    1. In 56%56\% of evaluated EntityQuestions instances across Llama-3-8B-Instruct, Mistral-7B-Instruct, and Gemma-2-9B-Instruct, drawing 1,000 independent samples per question at temperature T=1.0T=1.0 produced zero correct answers.
    2. When the ungenerated gold answer aGa_G was manually appended to the candidate set A~(o)\tilde{A}(o), the internal probing classifier ranked aGa_G higher than every sampled incorrect candidate (K∗=1K^* = 1) in 9%9\% of all questions, whereas the autoregressive generation probability P(a∣q)P(a \mid q) never improved K∗K^*.
    3. On average across all evaluated models, in 7.2%7.2\% of all test questions, the model simultaneously satisfied three conditions:
      • No correct answer was generated across 1,000 independent attempts (T=1.0T=1.0).
      • The token-level generation likelihood assigned to the gold answer was negligible: P(aG∣q)<0.01P(a_G \mid q) < 0.01.
      • The internal hidden-state probe achieved perfect ranking: K∗=1K^* = 1.

    This demonstrates that autoregressive token-level decoding mechanisms can completely suppress access to factual knowledge that is robustly encoded in the model's internal parameter representations.

  5. Knowl 5 — Knowledge-Aware Probing for Factual Correctness

    model/method

    Standard probing classifiers for QA factuality risk learning whether a model feels uncertain about a question rather than whether a specific answer aa is correct, because sampling random (question, answer) pairs typically pairs known questions with correct answers and unknown questions with incorrect answers. Knowledge-aware probing isolates representation of answer correctness from question-level familiarity.

    The training data collection procedure operates as follows:

    1. Filter for questions qq from the training split where the model MM generates the exact-match correct gold answer aGa_G under greedy decoding (ensuring MM knows the answer to qq).
    2. Use the correct greedy answer as the positive example (q,acorrect)(q, a_{\text{correct}}).
    3. Induce a hallucinated candidate by sampling additional outputs from MM at high temperature (T=2.0T=2.0) until an incorrect answer aincorrecta_{\text{incorrect}} is generated, and assign it as the negative example (q,aincorrect)(q, a_{\text{incorrect}}) for that same question.
    4. Extract the hidden state representation hM(q,a)h_M(q, a) from a specific layer of MM when prompted with (q,a)(q, a).
    5. Train a linear classifier using logistic regression to predict binary correctness from hM(q,a)h_M(q, a), ensuring zero subject or object entity overlap between training and test sets.
  6. Knowl 6 — Closed-Book QA Accuracy via Repeated Sampling and Internal Reranking

    data/table

    Increasing test-time compute by sampling 1,000 answer candidates (T=1.0T=1.0) and selecting the top candidate using the internal probe achieves a 12.1%12.1\% average relative accuracy improvement over greedy decoding across Llama-3-8B, Mistral-7B, and Gemma-2-9B. However, when the gold answer is included in the candidate pool whenever it was not sampled (Probe w. gold), performance increases by +52.7%+52.7\% relative to greedy, demonstrating that decoding failure locks away an additional ∼40%\sim 40\% relative gain that internal representations can reliably identify.

    Scoring Method Llama-3-8B Mistral-7B Gemma-2-9B Average
    Greedy 22.1 18.8 22.7 21.2
    Random 16.4* (-25.8%) 12.3* (-34.6%) 11.3* (-50.2%) 13.3* (-36.9%)
    Majority 23.7* (+7.2%) 19.7* (+4.8%) 22.6 (-0.4%) 22.0* (+3.9%)
    P(a∣q)P(a \mid q) 23.6* (+6.8%) 20.0* (+6.4%) 23.2* (+2.2%) 22.3* (+5.1%)
    Probe 25.4* (+14.9%) 22.0* (+17.0%) 23.7* (+4.4%) 23.7* (+12.1%)
    Oracle 44.2* (+100.0%) 38.9* (+106.9%) 49.8* (+119.4%) 44.3* (+108.8%)
    Probe w. gold 34.5* (+56.1%) 33.9* (+80.3%) 27.6* (+21.6%) 32.0* (+52.7%)

    Notes:

    • Accuracy values represent percentages over test splits across all four relations (P26, P264, P176, P50).
    • * denotes statistically significant differences (p<0.05p < 0.05) compared to greedy decoding.
    • Oracle assigns score 1 to correct sampled candidates and 0 to incorrect ones.
    • Probe w. gold manually inserts the ground-truth answer into the 1,000 sampled candidates when the generator fails to produce it.
  7. Knowl 7 — Closed-Form Probability of Success Under Repeated Sampling Inference Scaling

    theoretical result

    Let p=Pr⁡a∼M[a∈A(o)]p = \Pr_{a \sim M}[a \in A(o)] be the probability that a single sample generated by model MM is correct for factual triplet (s,r,o)(s, r, o). When drawing nn independent and identically distributed candidate answers from MM and selecting the candidate ranked highest by scoring function SMS_M, the probability that the selected top candidate is correct is given by:

    Pr⁡[Success((s,r,o);n,p,SM)]=∑i=0n(ni)pi(1−p)n−i[1−(1−K(s,r,o;SM)n−i)i]\Pr[\text{Success}((s, r, o); n, p, S_M)] = \sum_{i=0}^n \binom{n}{i} p^i (1 - p)^{n-i} \left[ 1 - \left( 1 - K(s, r, o; S_M)^{n-i} \right)^i \right]

    where:

    • (ni)pi(1−p)n−i\binom{n}{i} p^i (1 - p)^{n-i} is the probability of drawing exactly ii correct candidates and n−in - i incorrect candidates in the sample of size nn.
    • K(s,r,o;SM)K(s, r, o; S_M) is the pairwise ranking knowledge score representing the probability that a randomly chosen correct answer is scored higher than a randomly chosen incorrect answer by SMS_M.
    • K(s,r,o;SM)n−iK(s, r, o; S_M)^{n-i} is the probability that a single specific correct answer outranks all n−in - i incorrect candidates.
    • (1−K(s,r,o;SM)n−i)i\left( 1 - K(s, r, o; S_M)^{n-i} \right)^i is the probability that all ii correct candidates fail to outrank all n−in - i incorrect candidates.

    While an increase in K(s,r,o;SM)K(s, r, o; S_M) strictly increases the theoretical success probability Pr⁡[Success]\Pr[\text{Success}], KK scores reward ranking accuracy across all valid paraphrases of an answer, whereas inference scaling requires only a single correct candidate to outrank all incorrect candidates.

  8. Knowl 8 — Knowledge Metric with Plausibility Sanity Check

    definition

    To handle cases where a scoring function SMS_M might score an arbitrary, ungrammatical, or nonsensical string higher than a correct answer, the knowledge metric extends candidate evaluation over the full token sequence space A~M=V∗\tilde{A}_M = \mathcal{V}^*, where V\mathcal{V} is the model tokenizer vocabulary.

    A plausibility sanity-check indicator γ(q;SM)\gamma(q; S_M) is defined as:

    γ(q;SM)=I(∀a∈A~(o),a^∈A~M∖A~(o):SM(q,a)>SM(q,a^))\gamma(q; S_M) = \mathbb{I}\left( \forall a \in \tilde{A}(o), \hat{a} \in \tilde{A}_M \setminus \tilde{A}(o) : S_M(q, a) > S_M(q, \hat{a}) \right)

    where A~(o)\tilde{A}(o) is the set of semantically plausible entities of the required entity type.

    The extended per-question knowledge score is:

    Kq(s,r,o;SM)=γ(q;SM)⋅1∣Ω(s,r,o)∣∑(a,a~)∈Ω(s,r,o)I(SM(q,a)>SM(q,a~))K_q(s, r, o; S_M) = \gamma(q; S_M) \cdot \frac{1}{|\Omega(s, r, o)|} \sum_{(a, \tilde{a}) \in \Omega(s, r, o)} \mathbb{I}(S_M(q, a) > S_M(q, \tilde{a}))

    where Ω(s,r,o)=A(o)×(A~(o)∖A(o))\Omega(s, r, o) = A(o) \times (\tilde{A}(o) \setminus A(o)) is the set of valid-invalid plausible answer pairs.

  9. Knowl 9 — Empirical Disproof of Memorization in Hidden-State Probing Classifiers

    data/table

    To confirm that the internal probe measures authentic internal factual truthfulness rather than memorizing surface-level text or entity patterns from training pairs (q,a)(q, a), classifiers were evaluated on identical balanced test splits using input text features (TF-IDF), input token embedding mean-pooling (EMBED_MEAN), and model hidden representations h(f(q,a))h(f(q, a)) from Qwen3-32B.

    Input Representation Model Train Acc (%) Test Acc (%)
    - Random - 50.0
    TFIDF(f(q, a)) Logistic Regression 95.6 50.5
    TFIDF(f(q, a)) MLP (256, 256) 99.9 50.7
    TFIDF(f(q, a)) MLP (512, 512, 256, 128) 99.9 50.2
    EMBED_MEAN(f(q, a)) Logistic Regression 98.8 49.8
    EMBED_MEAN(f(q, a)) MLP (256, 256) 99.9 49.1
    EMBED_MEAN(f(q, a)) MLP (512, 512, 256, 128) 99.9 50.9
    h(f(q, a)) Logistic Regression 99.9 64.0

    Both surface lexical features (TF-IDF) and static token embedding pooled representations achieve near-perfect training accuracy (95.6%95.6\%–99.9%99.9\%) but collapse to chance performance on unseen test questions (49.1%49.1\%–50.9%50.9\%), proving that training set factual knowledge cannot generalize through surface text. In contrast, the probe trained on intermediate hidden representations h(f(q,a))h(f(q, a)) achieves 64.0%64.0\% test accuracy on balanced pairs.

  10. Knowl 10 — Layer-Wise Trajectory of Internal Factual Representations

    empirical result

    Layer-by-layer probing across the 32 transformer layers of Llama-3-8B-Instruct, Mistral-7B-Instruct, and Gemma-2-9B-Instruct reveals consistent structural dynamics in internal knowledge encoding:

    1. Knowledge scores KK are lowest in the initial layers and rise sharply through the lower third of the network.
    2. Knowledge scores plateau and stabilize in the upper two-thirds of the model, typically beginning around layers 11–12 of 32.
    3. A small decrease in KK score is frequently observed in the final 1–2 layers immediately preceding the output vocabulary projection head, indicating that representations directly driving autoregressive token selection exhibit degraded factuality ranking compared to late intermediate layers.

Coverage note — No substantial contributed material was omitted. All formal definitions, theoretical derivations, empirical findings across models, test-time scaling evaluations, layer-wise analyses, and validation controls are fully covered.

References

  1. 1.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  2. 2.Amos Azaria and Tom M. Mitchell. The internal state of an LLM knows when it’s lying. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pp. 967–976. Association for Computational Linguistics, 2023. doi: 10.18653 /V1/2023.FINDINGS-EMNLP.68. URL https://doi.org/10.18653/v1/2023.findings-emnlp.68.
  3. 3.Yonatan Belinkov and James R. Glass. Analysis methods in neural language processing: A survey. Trans. Assoc. Comput. Linguistics, 7:49–72, 2019. doi: 10.1162/TACL_A_00254. URL https://doi.org/10.1162/ tacl_a_00254.
  4. 4.Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. What do neural machine translation models learn about morphology? In Regina Barzilay and Min-Yen Kan (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 861–872, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1080. URL https://aclanthology.org/P17-1080/.
  5. 5.Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Re, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
  6. 6.Roger Brown and David McNeill. The “tip of the tongue” phenomenon. Journal of Verbal Learning and Verbal Behavior, 5(4):325–337, 1966. ISSN 0022-5371. doi: https://doi.org/10.1016/S0022-5371(66)80040-3. URL https://www.sciencedirect.com/science/article/pii/S0022537166800403.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https: //proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  8. 8.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=ETKGuby0hcs.
  9. 9.Noam Chomsky. Aspects of the Theory of Syntax. The MIT Press, 50 edition, 1965. ISBN 9780262527408. URL http://www.jstor.org/stable/j.ctt17kk81z.
  10. 10.Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. Crawling the internal knowledge-base of language models. In Andreas Vlachos and Isabelle Augenstein (eds.), Findings of the Association for Computational Linguistics: EACL 2023, pp. 1856–1869, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-eacl.139. URL https://aclanthology.org/2 023.findings-eacl.139.
  11. 11.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. Evaluating the ripple effects of knowledge editing in language models. Trans. Assoc. Comput. Linguistics, 12:283–298, 2024. doi: 10.1162/TACL_A_0 0644. URL https://doi.org/10.1162/tacl_a_00644.
  12. 12.Nicola De Cao, Wilker Aziz, and Ivan Titov. Editing factual knowledge in language models. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6491–6506, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021 .emnlp-main.522. URL https://aclanthology.org/2021.emnlp-main.522/.
  13. 13.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  14. 14.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schutze, and Yoav Goldberg. Measuring and improving consistency in pretrained language models. Trans. Assoc. Comput. Linguistics, 9:1012–1031, 2021. doi: 10.1162/TACL_A_00410. URL https://doi.org/10.1162/ta cl_a_00410.
  15. 15.Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. Probing for semantic evidence of composition by means of simple classification tasks. In Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, pp. 134–139, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/W16-2524. URL https://aclanthology.org/W16-2524/.
  16. 16.Tom Fawcett. An introduction to ROC analysis. Pattern Recognit. Lett., 27(8):861–874, 2006. doi: 10.1016/J.PA TREC.2005.10.010. URL https://doi.org/10.1016/j.patrec.2005.10.010.
  17. 17.Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau, and Anders Søgaard. Defining knowledge: Bridging epistemology and large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 16096–16111, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.900. URL https://aclanthology.org/2024.emnlp-main.900/.
  18. 18.Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. Does fine-tuning LLMs on new knowledge encourage hallucinations? In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 7765–7784, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.444. URL https://aclanthology.org/2024.emnlp-main.444/.
  19. 19.Daniela Gottesman and Mor Geva. Estimating knowledge in large language models without generating a single token. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 3994–4019, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.232. URL https://aclanthology.org/2024.emnlp-main.232/.
  20. 20.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  21. 21.J. A. Hanley and B. J. McNeil. The meaning and use of the area under a receiver operating characteristic (roc) curve. Radiology, 143(1):29–36, 1982. doi: 10.1148/radiology.143.1.7063747. URL https://pubmed.ncbi.nl m.nih.gov/7063747/.
  22. 22.Michael Hassid, Tal Remez, Jonas Gehring, Roy Schwartz, and Yossi Adi. The larger the better? improved LLM code-generation via budget reallocation. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=QJvfpWSpWm.
  23. 23.Audrey Huang, Adam Block, Dylan J Foster, Dhruv Rohatgi, Cyril Zhang, Max Simchowitz, Jordan T. Ash, and Akshay Krishnamurthy. Self-improvement in language models: The sharpening mechanism. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum ?id=WJaUkwci9o.
  24. 24.Dieuwke Hupkes, Sara Veldhoen, and Willem H. Zuidema. Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure. J. Artif. Intell. Res., 61:907–926, 2018. doi: 10.1613/JAIR.1.11196. URL https://doi.org/10.1613/jair.1.11196.
  25. 25.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  26. 26.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438, 2020. doi: 10.1162/tacl_a_00324. URL https://aclanthology.org/2020.tacl-1.28.
  27. 27.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  28. 28.Nora Kassner, Benno Krojer, and Hinrich Schutze. Are pretrained language models symbolic reasoners over knowledge? In Raquel Fernandez and Tal Linzen (eds.), Proceedings of the 24th Conference on Computational Natural Language Learning, pp. 552–564, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.conll-1.45. URL https://aclanthology.org/2020.conll-1.45/.
  29. 29.Nora Kassner, Oyvind Tafjord, Hinrich Schutze, and Peter Clark. BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 8849–8861, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.697. URL https://aclanthology.org/2021.emnlp-main.697/.
  30. 30.Kenneth Li, Oam Patel, Fernanda B. Viegas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023a. URL http://papers.nips.cc/paper_files/paper/2023/hash/81b8390 039b7302c909cb769f8b6cd93-Abstract-Conference.html.
  31. 31.Kenneth Li, Oam Patel, Fernanda B. Viegas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023b. URL http://papers.nips.cc/paper_files/paper/2023/hash/81b83 90039b7302c909cb769f8b6cd93-Abstract-Conference.html.
  32. 32.Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words. Trans. Mach. Learn. Res., 2022, 2022. URL https://openreview.net/forum?id=8s8K2UZGTZ.
  33. 33.Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness? In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 4791–4797, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.291. URL https://aclanthology.org/2023.emnlp-main.291/.
  34. 34.Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, 2024. URL https: //openreview.net/forum?id=aajyHYjjsk.
  35. 35.OpenAI. Introducing openai o1 preview, 2024. URL https://openai.com/index/introducing-openai-o1-p review/. https://openai.com/index/introducing-openai-o1-preview/.
  36. 36.Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov. Llms know more than they show: On the intrinsic representation of LLM hallucinations. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=KRnsX5Em3W.
  37. 37.Vaidehi Patil, Peter Hase, and Mohit Bansal. Can sensitive information be deleted from llms? objectives for defending against extraction attacks. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?i d=7erlRDoaV8.
  38. 38.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2463–2473, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1250. URL https://aclanthology.org/D19-1250.
  39. 39.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  40. 40.Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman. Weakly supervised detection of hallucinations in LLM activations. In Socially Responsible Language Modelling Research, 2023. URL https://openreview.net/forum?id=zNgdomlg4k.
  41. 41.Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15504–15522, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.828. URL https://aclanthology.org/2024.acl-long.828/.
  42. 42.Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5418–5426, 2020.
  43. 43.Juan Diego Rodriguez, Wenxuan Ding, Katrin Erk, and Greg Durrett. Rankalign: A ranking view of the generator-validator gap in large language models. arXiv preprint arXiv:2504.11381, 2025.
  44. 44.Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. Simple entity-centric questions challenge dense retrievers. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 6138–6148. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.EMNLP-MAIN.496. URL https://doi.org/10 .18653/v1/2021.emnlp-main.496.
  45. 45.Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov. Distinguishing ignorance from error in llm hallucinations. arXiv preprint arXiv:2410.22071, 2024.
  46. 46.Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge. Nature, 620(7972):172–180, 2023.
  47. 47.Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=4FWAwZtd2n.
  48. 48.Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. On early detection of hallucinations in factual question answering. In Ricardo Baeza-Yates and Francesco Bonchi (eds.), Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2024, Barcelona, Spain, August 25-29, 2024, pp. 2721–2732. ACM, 2024. doi: 10.1145/3637528.3671796. URL https://doi.org/10.1145/3637528. 3671796.
  49. 49.Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, and Yiqun Liu. Unsupervised real-time hallucination detection based on the internal states of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 14379–14391, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.854. URL https://aclanthology.org/2024.findings-acl.854/.
  50. 50.Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona. Confidence improves self-consistency in llms. arXiv preprint arXiv:2502.06233, 2025.
  51. 51.Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Leonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Rame, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024.
  52. 52.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pp. 5433–5442. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.EMNLP-MAIN.330. URL https://doi.org/10.18653/v1/2023.emnlp-main.330.
  53. 53.Rebecca Treiman, Charles Clifton Jr, Antje S Meyer, and Lee H Wurm. Language comprehension and production. Handbook of psychology, pp. 525–547, 2003.
  54. 54.Eduard Tulchinskii, Laida Kushnareva, Kristian Kuznetsov, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, and Serguei Barannikov. Listening to the wise few: Select-and-copy attention heads for multiple-choice qa. arXiv preprint arXiv:2410.02343, 2024.
  55. 55.Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36:74952–74965, 2023.
  56. 56.Denny Vrandecic and Markus Krotzsch. Wikidata: a free collaborative knowledgebase. Commun. ACM, 57 (10):78–85, sep 2014. ISSN 0001-0782. doi: 10.1145/2629489. URL https://doi.org/10.1145/2629489.
  57. 57.Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Bowen Ding, Zhikun Xu, Yidong Wang, Xiangkun Hu, Zheng Zhang, and Yue Zhang. Evaluating open-qa evaluation. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023a. URL http://papers.nips.cc/paper_files/paper/2023/hash/f323d59 4aa5d2c68154433a131c07959-Abstract-Datasets_and_Benchmarks.html.
  58. 58.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2023b. URL https://openreview.net/forum?i d=1PL1NIMMrw.
  59. 59.Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368, 2024.
  60. 60.An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024.
  61. 61.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  62. 62.Gal Yona, Roee Aharoni, and Mor Geva. Narrowing the knowledge evaluation gap: Open-domain question answering with multi-granularity answers. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6737–6751, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.365. URL https://aclanthology.org/2024.acl-long.365/.
  63. 63.Gal Yona, Or Honovich, Omer Levy, and Roee Aharoni. Keep guessing? when considering inference scaling, mind the baselines. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, pp. 5979–5991, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-8-89176-195-7. doi: 10.18653/v1/2025.findings-naacl.332. URL https://aclanthology.org/2025.findings-naacl.332/.
  64. 64.Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. Attention satisfies: A constraint-satisfaction lens on factual errors of language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=gfFVATffPd.
  65. 65.Shaolei Zhang, Tian Yu, and Yang Feng. Truthx: Alleviating hallucinations by editing large language models in truthful space. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 8908–8949. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL-LONG.483. URL https://doi.org/10.18653/v1/2024.acl-long.483.
  66. 66.Eric Zhao, Pranjal Awasthi, and Sreenivas Gollapudi. Sample, scrutinize and scale: Effective inference-time search by scaling verification. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=wl3eI4wiE5.
  67. 67.Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. Can we edit factual knowledge by in-context learning? In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pp. 4862–4876. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.EMNLP-MAIN.296. URL https://doi.org/10.18653/v1/2023.emnlp-main.296.
  68. 68.Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen. Mquake: Assessing knowledge editing in language models via multi-hop questions. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pp. 15686–15702. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.EMNLP-MAIN.971. URL https://doi.org/10.18653/v1/2023.emnlp-main.971.
  69. 69.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Gekhman, Z., et al. “Inside-Out: Hidden Factual Knowledge in LLMs”. arXiv, 2025, http://arxiv.org/abs/2503.15299v4.
APA
Gekhman, Z., David, E. B., Orgad, H., Ofek, E., Belinkov, Y., Szpektor, I., Herzig, J., & Reichart, R. (2025). Inside-Out: Hidden Factual Knowledge in LLMs. arXiv. http://arxiv.org/abs/2503.15299v4
Chicago
Gekhman, Z., E. B. David, H. Orgad, et al. 2025. “Inside-Out: Hidden Factual Knowledge in LLMs”. arXiv. http://arxiv.org/abs/2503.15299v4.
Harvard
Gekhman, Z. et al. (2025) “Inside-Out: Hidden Factual Knowledge in LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.15299v4.
Vancouver
1. Gekhman Z, David EB, Orgad H, Ofek E, Belinkov Y, Szpektor I, Herzig J, Reichart R (2025) Inside-Out: Hidden Factual Knowledge in LLMs. arXiv

BibTeX

@article{gekhman2025inside,
  title = {Inside-Out: Hidden Factual Knowledge in LLMs},
  author = {Gekhman, Zorik and David, Eyal Ben and Orgad, Hadas and Ofek, Eran and Belinkov, Yonatan and Szpektor, Idan and Herzig, Jonathan and Reichart, Roi},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.15299v4},
  eprint = {2503.15299}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/