LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Hadas OrgadMichael TokerZorik GekhmanRoi ReichartIdan SzpektorHadas KotekYonatan Belinkov

article2025ICLR298 citations

Demonstrates that large language models often internally represent the correct answer while generating hallucinations, and identifies token-level truthfulness signals that predict specific error types despite failing to transfer across datasets.

Listen

Large language models frequently generate inaccurate or biased information, widely known as hallucinations. While most conventional methods evaluate these errors based on external user perception or simple output confidence, this human-centric perspective fails to capture how models internally encode truthfulness. Understanding how models internally represent errors is critical for building trustworthy systems and implementing precise safety mechanisms. The article evaluates how truthfulness is represented across the internal layers and tokens of language models, demonstrating how these internal signals can be used to detect errors, classify error types, and select correct answers.

To conduct this evaluation, the researchers tested four open-source language models across ten diverse datasets covering factual retrieval, commonsense reasoning, natural language inference, sentiment analysis, and mathematics. They allowed models to generate unrestricted, long-form answers to mirror practical use. The analysis involved training lightweight linear classifiers, called probing classifiers, on intermediate model activations across different layers and token positions. The study compared these probing methods against common baseline techniques, such as output probability aggregation and direct self-evaluation prompting, and evaluated their performance across tasks, error distributions, and multiple generated samples.

Key findings reveal that internal truthfulness signals are highly concentrated in the exact answer tokens rather than the prompt or final generated tokens. Exploiting activations at these exact answer tokens substantially improved error detection accuracy, outperforming all logit- and probability-based baselines across all datasets. Second, the article demonstrates that truthfulness encoding is skill-specific rather than universal. Probes trained on one dataset failed to generalize to fundamentally different tasks beyond what simple output probabilities could already predict, refuting prior assumptions of a single, universal internal truth mechanism. Third, internal representations successfully predicted fine-grained error types, such as whether a model would consistently repeat the same error or generate widespread guesses. Finally, the analysis revealed a severe discrepancy between internal knowledge and external output: even when models consistently generated incorrect responses, their internal representations often identified the correct answer among resampled candidates, allowing probe-based selection to improve task accuracy by 30 to 40 percentage points on difficult error categories.

These findings imply that relying on a single, general-purpose truth detector introduces operational risks, as models handle different reasoning skills through distinct internal mechanisms. However, organizations can build highly reliable, task-specific error detectors for high-stakes domains such as medicine and law by targeting exact answer tokens. The results also show that models often possess the required factual knowledge internally but fail to generate it due to decoding mechanisms that prioritize text likelihood over truthfulness, indicating significant potential to recover accurate answers without retraining entire models.

Organizations deploying language models should avoid universal truth-detecting filters and instead implement task-specific probing architectures focused on exact answer tokens within specialized workflows. For complex domains, engineering teams can combine internal probing with answer resampling or targeted retrieval interventions based on the predicted error type. Further development should focus on building lightweight extraction modules for real-time answer token identification and testing internal decoding interventions on open-ended generation tasks.

The conclusions are limited to open-source models, as the methodology requires direct access to intermediate neural network activations, making it inapplicable to closed, black-box systems. Additionally, the benchmarks focused on structured tasks with verifiable reference labels rather than open-ended creative text. Confidence in these findings is high across standard question-answering and reasoning benchmarks, though caution is warranted when attempting to extrapolate probe behavior to unseen task domains without task-specific validation.

arXiv: 2410.02707
Cover for LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Abstract

Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states encode information regarding the truthfulness of their outputs, and that this information can be utilized to detect errors. In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. Yet, we show that such error detectors fail to generalize across datasets, implying that -- contrary to prior claims -- truthfulness encoding is not universal but rather multifaceted. Next, we show that internal representations can also be used for predicting the types of errors the model is likely to make, facilitating the development of tailored mitigation strategies. Lastly, we reveal a discrepancy between LLMs' internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one. Taken together, these insights deepen our understanding of LLM errors from the model's internal perspective, which can guide future research on enhancing error analysis and mitigation.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Better Error Detection
  • 3.1 Task Definition
  • 3.2 Experimental Setup
  • 3.3 Results
  • 4 Generalization Between Tasks
  • 5 Investigating Error Types
  • 5.1 Taxonomy of Errors
  • 5.2 Predicting Error Types
  • 6 Detecting the Correct Answer
  • 7 Discussion and Conclusions
  • 8 Reproducibility Statement
  • References
  • A Implementation Details
  • A.1 Task Specific Error Detection
  • A.2 Probing: Implementation Details
  • A.3 Datasets
  • A.4 Baselines: Implementation Details
  • B Full Error Detection Results
  • C Full Generalization Results
  • D Taxonomy of Errors
  • D.1 Error Taxonomy Design Choices
  • D.2 Results on Additional Datasets
  • D.3 Qualitative Examples
  • E Detecting the Correct Answer Full Results
  • F Practical Guidance on Integrating Insights from this Paper into Model Development Workflows

Knowls

  1. Knowl 1 — Exact Answer Token Probing for LLM Error Detection

    model/method

    To detect whether a large language model (LLM) generated response y^\hat{y} to a prompt pp is correct or incorrect without external tools, intermediate representations at specific "exact answer tokens" are analyzed. Exact answer tokens are defined as the minimal substring in a long-form generated response whose alteration changes the factual correctness of the answer, disregarding surrounding conversational filler or reasoning text.

    Given an input prompt p=(q1,…,qm)p = (q_1, \dots, q_m) and a generated sequence y^=(y^1,…,y^n)\hat{y} = (\hat{y}_1, \dots, \hat{y}_n), the exact answer token span is identified as indices [tstart,tend]⊆{1,…,n}[t_{\text{start}}, t_{\text{end}}] \subseteq \{1, \dots, n\}. For each transformer layer l∈{1,…,L}l \in \{1, \dots, L\}, the activation vector hl,tend∈Rdh_{l, t_{\text{end}}} \in \mathbb{R}^d is extracted from the output of the final Multi-Layer Perceptron (MLP) at the last token of the exact answer span.

    A linear probing classifier (specifically an L2L_2-regularized logistic regression model solved via L-BFGS) is trained on the activation vectors hl,tendh_{l, t_{\text{end}}} across training instances to output a predicted correctness probability P(correct∣hl,tend)P(\text{correct} \mid h_{l, t_{\text{end}}}). The layer index ll is selected based on maximum validation performance as measured by the area under the ROC curve (AUC).

  2. Knowl 2 — Layer and Token Trajectory of Truthfulness Representations in LLMs

    empirical result

    Systematic linear probing across all transformer layers l∈{1,…,L}l \in \{1, \dots, L\} and token positions tt from the final question token through the entire generated response in autoregressive models (including Mistral-7B, Mistral-7B-Instruct, Llama-3-8B, and Llama-3-8B-Instruct) reveals a structured trajectory of truthfulness signals:

    1. Token Dynamics:

      • A pronounced initial truthfulness signal appears immediately at the final prompt/question token, indicating that the hidden state already encodes the model's general capability to answer the question.
      • During the generation of introductory tokens and filler phrases, this signal degrades.
      • The truthfulness signal peaks sharply at the exact answer tokens (specifically the last token of the exact answer span), where the area under the ROC curve (AUC) achieves its global maximum.
      • After the exact answer tokens, the signal drops again, followed by a slight rise at the final generation token (</s></\text{s}>).
    2. Layer Dynamics:

      • Truthfulness encoding is weak in the earliest transformer layers, rises steadily, and peaks in the middle-to-later layers before slightly decreasing in the final output layer, consistent with the localized emergence of factual association representations in transformer architectures.
  3. Knowl 3 — Error Detection Performance of Exact Answer Probing Across Datasets

    data/table

    Error detection effectiveness is evaluated using the Area Under the Receiver Operating Characteristic Curve (AUC-ROC) across open-ended question answering (TriviaQA), gender-bias coreference resolution (Winobias), and arithmetic reasoning (Math). The table compares token-probability baselines (Logits-mean, Logits-min), prompt-based self-evaluation (P(True)P(\text{True})), static-token probing (last generated token [−1][-1], second to last [−2][-2], and end-of-question token), and probing at the exact answer token:

    Method Mistral-7B-Instruct Llama-3-8B-Instruct
    TriviaQA Winobias Math TriviaQA Winobias Math
    Logits-mean 0.60±0.0090.60 \pm 0.009 0.56±0.0170.56 \pm 0.017 0.55±0.0290.55 \pm 0.029 0.66±0.0050.66 \pm 0.005 0.60±0.0260.60 \pm 0.026 0.75±0.0180.75 \pm 0.018
    Logits-mean-exact 0.68±0.0070.68 \pm 0.007 0.54±0.0120.54 \pm 0.012 0.51±0.0050.51 \pm 0.005 0.71±0.0060.71 \pm 0.006 0.55±0.0190.55 \pm 0.019 0.80±0.0210.80 \pm 0.021
    Logits-min 0.63±0.0080.63 \pm 0.008 0.59±0.0120.59 \pm 0.012 0.51±0.0170.51 \pm 0.017 0.74±0.0070.74 \pm 0.007 0.61±0.0240.61 \pm 0.024 0.75±0.0160.75 \pm 0.016
    Logits-min-exact 0.75±0.0060.75 \pm 0.006 0.53±0.0130.53 \pm 0.013 0.71±0.0090.71 \pm 0.009 0.79±0.0060.79 \pm 0.006 0.61±0.0190.61 \pm 0.019 0.89±0.0180.89 \pm 0.018
    P(True)P(\text{True}) 0.66±0.0060.66 \pm 0.006 0.45±0.0210.45 \pm 0.021 0.48±0.0220.48 \pm 0.022 0.73±0.0080.73 \pm 0.008 0.59±0.0200.59 \pm 0.020 0.62±0.0170.62 \pm 0.017
    P(True)P(\text{True})-exact 0.74±0.0030.74 \pm 0.003 0.40±0.0210.40 \pm 0.021 0.60±0.0250.60 \pm 0.025 0.73±0.0050.73 \pm 0.005 0.63±0.0140.63 \pm 0.014 0.59±0.0180.59 \pm 0.018
    Probe @ Last generated [−1][-1] 0.71±0.0060.71 \pm 0.006 0.82±0.0040.82 \pm 0.004 0.74±0.0080.74 \pm 0.008 0.81±0.0050.81 \pm 0.005 0.86±0.0070.86 \pm 0.007 0.82±0.0160.82 \pm 0.016
    Probe @ Before last [−2][-2] 0.73±0.0040.73 \pm 0.004 0.85±0.0040.85 \pm 0.004 0.74±0.0070.74 \pm 0.007 0.75±0.0050.75 \pm 0.005 0.88±0.0050.88 \pm 0.005 0.79±0.0200.79 \pm 0.020
    Probe @ End of question 0.76±0.0080.76 \pm 0.008 0.82±0.0110.82 \pm 0.011 0.72±0.0070.72 \pm 0.007 0.77±0.0070.77 \pm 0.007 0.80±0.0180.80 \pm 0.018 0.72±0.0230.72 \pm 0.023
    Probe @ Exact answer 0.85±0.004\mathbf{0.85 \pm 0.004} 0.92±0.005\mathbf{0.92 \pm 0.005} 0.92±0.008\mathbf{0.92 \pm 0.008} 0.83±0.002\mathbf{0.83 \pm 0.002} 0.93±0.004\mathbf{0.93 \pm 0.004} 0.95±0.027\mathbf{0.95 \pm 0.027}

    Extracting representations at the exact answer token consistently yields the highest AUC across all tested architectures and datasets, significantly outperforming output probability metrics, self-prompted evaluations, and static token locations.

  4. Knowl 4 — Skill-Specificity and Non-Universality of Internal Truthfulness Representations

    empirical result

    When linear probing classifiers trained on exact answer activations from one dataset are evaluated across 10 diverse tasks (TriviaQA, HotpotQA with/without context, Natural Questions with context, Movies, Winobias, Winogrande, MNLI, Math, IMDB), cross-task generalization exhibits the following properties:

    1. Apparent vs. Baseline Transfer: Evaluating cross-task transfer with absolute AUC, defined as max⁡(AUC,1−AUC)\max(\text{AUC}, 1 - \text{AUC}), initially yields scores above 0.50.5 across most dataset pairs. However, when subtracting the performance of the strongest logit-based baseline (Logits-min-exact\text{Logits-min-exact}), the probe's transfer margin drops to zero or becomes negative for out-of-domain task pairs. This indicates that cross-task generalization above chance is largely driven by surface uncertainty features already present in the output logits rather than a universal internal truthfulness direction.

    2. Skill-Specific Clustering: Probing classifiers maintain positive transfer beyond logit baselines only between tasks requiring similar underlying cognitive skills:

      • Parametric factual retrieval: TriviaQA, HotpotQA without context, and Movies generalize mutually.
      • Common-sense and coreference reasoning: Winobias, Winogrande, and MNLI generalize mutually.

    These results demonstrate that large language models do not maintain a single monolithic or universal representation of truth, but rather multiple multifaceted, skill-specific truthfulness mechanisms.

  5. Knowl 5 — Behavioral Error Taxonomy from Repeated Generation Distributions

    definition

    A behavioral taxonomy classifies large language model generation errors by sampling K=30K = 30 responses at temperature T=1.0T = 1.0 for each question and characterizing the resulting empirical distribution. Let NuniqueN_{\text{unique}} be the number of unique answers generated, fcorrectf_{\text{correct}} be the frequency of the correct ground-truth answer (0≤fcorrect≤300 \le f_{\text{correct}} \le 30), and fmode, wrongf_{\text{mode, wrong}} be the frequency of the most frequent incorrect answer. The error classes are defined as:

    • (A) Refuses to answer: The model explicitly states it cannot answer in at least half of the samples (frefuse≥15f_{\text{refuse}} \ge 15).
    • (B) Consistently correct: The model produces the correct answer in at least half of the samples (fcorrect≥15f_{\text{correct}} \ge 15).
      • (B1) Always correct: fcorrect=30f_{\text{correct}} = 30.
      • (B2) Mostly correct: 15≤fcorrect<3015 \le f_{\text{correct}} < 30, with occasional errors.
    • (C) Consistently incorrect: The model generates the same incorrect answer in at least half of the samples (fmode, wrong≥15f_{\text{mode, wrong}} \ge 15).
      • (C1) Correct never produced: fcorrect=0f_{\text{correct}} = 0.
      • (C2) Correct appears at least once: 1≤fcorrect<151 \le f_{\text{correct}} < 15.
    • (D) Two competing answers: Both the correct answer and a specific incorrect answer appear at similar rates, defined as ∣fcorrect−fmode, wrong∣≤5|f_{\text{correct}} - f_{\text{mode, wrong}}| \le 5, with fcorrect≥5f_{\text{correct}} \ge 5 and fmode, wrong≥5f_{\text{mode, wrong}} \ge 5.
    • (E) Many answers: High dispersion with Nunique>10N_{\text{unique}} > 10.
      • (E1) Non-correct: fcorrect=0f_{\text{correct}} = 0.
      • (E2) Correct appears: fcorrect≥1f_{\text{correct}} \ge 1.

    This non-orthogonal taxonomy covers 96%96\% of errors observed on TriviaQA for Mistral-7B-Instruct.

  6. Knowl 6 — Intrinsic Predictability of Behavioral Error Types from Intermediate States

    data/table

    One-vs-all logistic regression probes trained on intermediate activations of greedy decoding responses (at the last exact answer token) can predict the multi-sample behavioral error type (A through E) of an input question. The table reports AUC scores for classifying error types on TriviaQA across four models:

    Error Type Mistral-7B Mistral-7B-Instruct Llama-3-8B Llama-3-8B-Instruct
    (A) Refuses to answer 0.86±0.0020.86 \pm 0.002 0.85±0.0110.85 \pm 0.011 0.87±0.0020.87 \pm 0.002 0.88±0.0140.88 \pm 0.014
    (B) Consistently correct 0.88±0.0010.88 \pm 0.001 0.82±0.0080.82 \pm 0.008 0.86±0.0010.86 \pm 0.001 0.81±0.0020.81 \pm 0.002
    (C) Consistently incorrect 0.59±0.0020.59 \pm 0.002 0.67±0.0020.67 \pm 0.002 0.59±0.0020.59 \pm 0.002 0.64±0.0030.64 \pm 0.003
    (D) Two competing answers 0.63±0.0020.63 \pm 0.002 0.68±0.0060.68 \pm 0.006 0.61±0.0010.61 \pm 0.001 0.65±0.0040.65 \pm 0.004
    (E) Many answers 0.90±0.0010.90 \pm 0.001 0.84±0.0030.84 \pm 0.003 0.89±0.0010.89 \pm 0.001 0.89±0.0010.89 \pm 0.001

    The predictability of refusal (A), consistent correctness (B), and high dispersion (E) demonstrates that the intermediate activations of a single greedy response contain structured representations of the model's epistemic state and distributional behavior across repeated samplings.

  7. Knowl 7 — Internal Truthfulness Disconnect and Probe-Based Resample Selection

    data/table

    When an LLM generates K=30K = 30 candidate responses (T=1.0T = 1.0), selecting the response with the highest predicted correctness probability from an exact-answer error-detection probe significantly outperforms greedy decoding, random candidate sampling, and majority voting. This improvement is concentrated in questions where the model outwardly exhibits low or zero preference for the correct answer during text generation.

    The following table presents accuracy across error types for Mistral-7B-Instruct on TriviaQA and Math:

    Error Subtype TriviaQA Math
    Greedy Random Majority Probing Greedy Random Majority Probing
    All 0.630.63 0.640.64 0.670.67 0.71\mathbf{0.71} 0.550.55 0.520.52 0.570.57 0.70\mathbf{0.70}
    (B2) Mostly correct 0.880.88 0.830.83 0.99\mathbf{0.99} 0.890.89 0.870.87 0.840.84 1.00\mathbf{1.00} 0.960.96
    (C2) Mostly incorrect, ≥1\ge 1 correct 0.110.11 0.150.15 0.000.00 0.53\mathbf{0.53} 0.100.10 0.200.20 0.000.00 0.82\mathbf{0.82}
    (D) Two competing answers 0.320.32 0.450.45 0.500.50 0.78\mathbf{0.78} — — — —
    (E2) Many answers, ≥1\ge 1 correct 0.230.23 0.190.19 0.380.38 0.56\mathbf{0.56} — — — —

    In category (C2), where greedy decoding achieves only 10%10\%--11%11\% accuracy and majority voting drops to 0%0\%, probe-based selection achieves 53%53\% accuracy on TriviaQA and 82%82\% on Math. In category (D), accuracy increases from 32%32\% (greedy) and 50%50\% (majority) to 78%78\%. This indicates a fundamental misalignment: autoregressive generation mechanisms prioritize high-likelihood surface tokens over truthfulness, even when internal representations accurately distinguish the true answer among candidate generations.

  8. Knowl 8 — Exact Answer Token Extraction Algorithm for Free-Form Generations

    algorithm

    To extract the exact answer token span [tstart,tend][t_{\text{start}}, t_{\text{end}}] from an unconstrained long-form response y^\hat{y} given a question qq and reference answer yy, an automated pipeline combines substring search heuristics and instruction-tuned LLM extraction.

    Input: Question string qq, Generated long-form response y^=(w1,…,wn)\hat{y} = (w_1, \dots, w_n), Reference answer yy
    Output: Token span [tstart,tend][t_{\text{start}}, t_{\text{end}}] of the exact answer, or null if no answer
    if task has a closed, finite label set then
        Search for first valid label match in y^\hat{y}
        if match found then
            return token span of the matched label
        end if
    end if
    if y^\hat{y} contains reference answer yy as a substring then
        Identify token span [tstart,tend][t_{\text{start}}, t_{\text{end}}] corresponding to the first occurrence of yy in y^\hat{y}
        return [tstart,tend][t_{\text{start}}, t_{\text{end}}]
    end if
    Prompt an instruction-tuned LLM (e.g., Mistral-7B-Instruct) with few-shot exemplars:
    "Extract from the following long answer the short answer, only the relevant tokens. If the long answer does not answer the question, output NO ANSWER."
    for attempt=1\text{attempt} = 1 to 55 do
        sextracted←LLM_extract(q,y^)s_{\text{extracted}} \leftarrow \text{LLM\_extract}(q, \hat{y})
        if sextracted="NO ANSWER"s_{\text{extracted}} = \text{"NO ANSWER"} then
            return null
        end if
        if sextracteds_{\text{extracted}} is a contiguous substring of y^\hat{y} then
            Locate token span [tstart,tend][t_{\text{start}}, t_{\text{end}}] corresponding to sextracteds_{\text{extracted}} in y^\hat{y}
            return [tstart,tend][t_{\text{start}}, t_{\text{end}}]
        end if
    end for
    return null

    Questions where extraction returns null or fails to match a valid substring within 5 attempts are excluded to prevent correlation between extraction artifacts and generation errors.

Coverage note — Omitted full appendix tables containing raw AUC metrics for all 10 datasets across all four models, as they confirm the exact same empirical trends captured in the representative summary tables.

References

  1. 1.Alexandre Allauzen. Error detection in confusion network. In 8th Annual Conference of the International Speech Communication Association, INTERSPEECH 2007, Antwerp, Belgium, August 27-31, 2007, pp. 1749–1752. ISCA, 2007. doi: 10.21437/INTERSPEECH.2007-490. URL https://doi.org/10.21437/Interspeech.2007-490.
  2. 2.Amos Azaria and Tom Mitchell. The internal state of an llm knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 967–976, 2023.
  3. 3.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023.
  4. 4.Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances, 2021. URL https://arxiv.org/abs/2102.12452.
  5. 5.Samuel J. Bell, Helen Yannakoudakis, and Marek Rei. Context is key: Grammatical error detection with contextual word representations. In Helen Yannakoudakis, Ekaterina Kochmar, Claudia Leacock, Nitin Madnani, Ildikó Pilán, and Torsten Zesch (eds.), Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, BEA@ACL 2019, Florence, Italy, August 2, 2019, pp. 103–115. Association for Computational Linguistics, 2019. doi: 10.18653/V1/W19-4410. URL https://doi.org/10.18653/v1/w19-4410.
  6. 6.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. On identifiability in transformers. In 8th International Conference on Learning Representations (ICLR 2020)(virtual). International Conference on Learning Representations, 2020.
  8. 8.Lennart Bürger, Fred A Hamprecht, and Boaz Nadler. Truth is universal: Robust detection of lies in llms. arXiv preprint arXiv:2407.12831, 2024.
  9. 9.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827, 2022.
  10. 10.Andrew Caines, Christian Bentz, Kate M. Knill, Marek Rei, and Paula Buttery. Grammatical error detection in transcriptions of spoken english. In Donia Scott, Nuria Bel, and Chengqing Zong (eds.), Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pp. 2144–2162. International Committee on Computational Linguistics, 2020. doi: 10.18653/V1/2020.COLING-MAIN.195. URL https://doi.org/10.18653/v1/2020.coling-main.195.
  11. 11.Sky CH-Wang, Benjamin Van Durme, Jason Eisner, and Chris Kedzie. Do androids know they’re only dreaming of electric sheep?, 2023.
  12. 12.Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. INSIDE: LLMs’ internal states retain the power of hallucination detection. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Zj12nzlQbz.
  13. 13.Wei Chen, Sankaranarayanan Ananthakrishnan, Rohit Kumar, Rohit Prasad, and Prem Natarajan. ASR error detection in a conversational spoken language translation system. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2013, Vancouver, BC, Canada, May 26-31, 2013, pp. 7418–7422. IEEE, 2013. doi: 10.1109/ICASSP.2013.6639104. URL https://doi.org/10.1109/ICASSP.2013.6639104.
  14. 14.Yong Cheng and Mofan Duan. Chinese grammatical error detection based on BERT model. In Erhong YANG, Endong XUN, Baolin ZHANG, and Gaoqi RAO (eds.), Proceedings of the 6th Workshop on Natural Language Processing Techniques for Educational Applications, pp. 108–113, Suzhou, China, December 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.nlptea-1.15.
  15. 15.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. Dola: Decoding by contrasting layers improves factuality in large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Th6NyL07na.
  16. 16.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12, 2021.
  17. 17.Rahhal Errattahi, Asmaa El Hannani, and Hassan Ouahmane. Automatic speech recognition errors detection and correction: A review. In Mourad Abbas and Ahmed Abdelali (eds.), 1st International Conference on Natural Language and Speech Processing, ICNLSP 2015, Algiers, Algeria, October 18-19, 2015, volume 128 of Procedia Computer Science, pp. 32–37. Elsevier, 2015. doi: 10.1016/J.PROCS.2018.03.005. URL https://doi.org/10.1016/j.procs.2018.03.005.
  18. 18.Dan Flickinger, Michael Wayne Goodman, and Woodley Packard. Uw-stanford system description for AESW 2016 shared task on grammatical error detection. In Joel R. Tetreault, Jill Burstein, Claudia Leacock, and Helen Yannakoudakis (eds.), Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications, BEA@NAACL-HLT 2016, June 16, 2016, San Diego, California, USA, pp. 105–111. The Association for Computer Linguistics, 2016. doi: 10.18653/V1/W16-0511. URL https://doi.org/10.18653/v1/w16-0511.
  19. 19.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. Rarr: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 16477–16508, 2023.
  20. 20.Zorik Gekhman, Roee Aharoni, Genady Beryozkin, Markus Freitag, and Wolfgang Macherey. KoBE: Knowledge-based machine translation evaluation. In Trevor Cohn, Yulan He, and Yang Liu (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3200–3207, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.287. URL https://aclanthology.org/2020.findings-emnlp.287.
  21. 21.Zorik Gekhman, Dina Zverinski, Jonathan Mallinson, and Genady Beryozkin. RED-ACE: Robust error detection for ASR using confidence embeddings. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (eds.), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 2800–2808, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.180. URL https://aclanthology.org/2022.emnlp-main.180.
  22. 22.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. TrueTeacher: Learning factual consistency evaluation with large language models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 2053–2070, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.127. URL https://aclanthology.org/2023.emnlp-main.127.
  23. 23.Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. Does fine-tuning llms on new knowledge encourage hallucinations?, 2024.
  24. 24.Zorik Gekhman, Eyal Ben David, Hadas Orgad, Eran Ofek, Yonatan Belinkov, Idan Szpector, Jonathan Herzig, and Roi Reichart. Inside-out: Hidden factual knowledge in llms. arXiv preprint arXiv:2503.15299, 2025.
  25. 25.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023.
  26. 26.Daniela Gottesman and Mor Geva. Estimating knowledge in large language models without generating a single token. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024), Miami, Florida, 2024. Association for Computational Linguistics.
  27. 27.Nuno M Guerreiro, Elena Voita, and André FT Martins. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 1059–1075, 2023.
  28. 28.Stevan Harnad. Language writ large: Llms, chatgpt, grounding, meaning and understanding. arXiv preprint arXiv:2402.02243, 2024.
  29. 29.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. q²: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7856–7870, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.619. URL https://aclanthology.org/2021.emnlp-main.619.
  30. 30.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. TRUE: Re-evaluating factual consistency evaluation. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (eds.), Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3905–3920, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.287. URL https://aclanthology.org/2022.naacl-main.287.
  31. 31.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232, 2023a.
  32. 32.Yuheng Huang, Jiayang Song, Zhijie Wang, Huaming Chen, and Lei Ma. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv preprint arXiv:2307.10236, 2023b.
  33. 33.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
  34. 34.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b, 2023. URL https://arxiv.org/abs/2310.06825.
  35. 35.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, 2017.
  36. 36.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  37. 37.Sudhanshu Kasewa, Pontus Stenetorp, and Sebastian Riedel. Wronging a right: Generating better errors to improve grammatical error detection. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 4977–4983. Association for Computational Linguistics, 2018. URL https://aclanthology.org/D18-1541/.
  38. 38.Hadas Kotek, Rikker Dockum, and David Sun. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, pp. 12–24, 2023.
  39. 39.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. Evaluating the factual consistency of abstractive text summarization. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9332–9346, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.750. URL https://aclanthology.org/2020.emnlp-main.750.
  40. 40.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=VD-AYtP0dve.
  41. 41.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics, 2019.
  42. 42.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177, 2022. doi: 10.1162/tacl_a_00453. URL https://aclanthology.org/2022.tacl-1.10.
  43. 43.Benjamin A Levinstein and Daniel A Herrmann. Still no lie detector for language models: Probing empirical and conceptual roadblocks. Philosophical Studies, pp. 1–27, 2024.
  44. 44.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33: 9459–9474, 2020.
  45. 45.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36, 2024.
  46. 46.Wei Li and Houfeng Wang. Detection-correction structure via general language model for grammatical error correction. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 1748–1763. Association for Computational Linguistics, 2024. URL https://aclanthology.org/2024.acl-long.96.
  47. 47.Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Yi Wang, Zhonghao Wang, Feiyu Xiong, et al. Internal consistency and self-feedback in large language models: A survey. arXiv preprint arXiv:2407.14507, 2024.
  48. 48.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  49. 49.Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness? In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 4791–4797, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.291. URL https://aclanthology.org/2023.emnlp-main.291.
  50. 50.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. A token-level reference-free hallucination detection benchmark for free-form text generation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6723–6737, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.464. URL https://aclanthology.org/2022.acl-long.464.
  51. 51.Chi-kiu Lo. YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources. In Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, André Martins, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Matt Post, Marco Turchi, and Karin Verspoor (eds.), Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pp. 507–513, Florence, Italy, August 2019. Association for Computational Linguistics. doi: 10.18653/v1/W19-5358. URL https://aclanthology.org/W19-5358.
  52. 52.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 142–150, Portland, Oregon, USA, June 2011. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/P11-1015.
  53. 53.Potsawee Manakul, Adian Liusie, and Mark Gales. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 9004–9017, 2023.
  54. 54.Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, 2023.
  55. 55.Alessia McGowan, Yunlai Gui, Matthew Dobbs, Sophia Shuster, Matthew Cotter, Alexandria Selloni, Marianne Goodman, Agrima Srivastava, Guillermo A Cecchi, and Cheryl M Corcoran. Chatgpt and bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Research, 326:115334, 2023.
  56. 56.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 36, 2022. arXiv:2202.05262.
  57. 57.Beren Millidge. LLMs confabulate not hallucinate. Beren’s Blog, March 2023. URL https://www.beren.io/2023-03-19-LLMs-confabulate-not-hallucinate/.
  58. 58.Ritika Mishra and Navjot Kaur. A survey of spelling error detection and correction techniques. International Journal of Computer Trends and Technology, 4(3):372–374, 2013.
  59. 59.nostalgebraist. Interpreting gpt: The logit lens. LessWrong blog post, 2020. URL https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens. Accessed: 2024-11-18.
  60. 60.Chris Olah, Nelson Elhage, Neel Nanda, Catherine Schubert, Daniel Filan, et al. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. URL https://transformer-circuits.pub/2023/monosemantic-features/index.html.
  61. 61.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  62. 62.Thomas Pellegrini and Isabel Trancoso. Error detection in broadcast news ASR using markov chains. In Zygmunt Vetulani (ed.), Human Language Technology. Challenges for Computer Science and Linguistics - 4th Language and Technology Conference, LTC 2009, Poznan, Poland, November 6-8, 2009, Revised Selected Papers, volume 6562 of Lecture Notes in Computer Science, pp. 59–69. Springer, 2009. doi: 10.1007/978-3-642-20095-3_6. URL https://doi.org/10.1007/978-3-642-20095-3_6.
  63. 63.Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. Learning compact metrics for MT. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 751–762, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.58. URL https://aclanthology.org/2021.emnlp-main.58.
  64. 64.Gaoqi Rao, Erhong Yang, and Baolin Zhang. Overview of NLPTEA-2020 shared task for Chinese grammatical error diagnosis. In Erhong YANG, Endong XUN, Baolin ZHANG, and Gaoqi RAO (eds.), Proceedings of the 6th Workshop on Natural Language Processing Techniques for Educational Applications, pp. 25–35, Suzhou, China, December 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.nlptea-1.4.
  65. 65.Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman. Weakly supervised detection of hallucinations in llm activations. arXiv preprint arXiv:2312.02798, 2023.
  66. 66.Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, SM Tonmoy, Aman Chadha, Amit P Sheth, and Amitava Das. The troubling emergence of hallucination in large language models–an extensive definition, quantification, and prescriptive remediations. arXiv preprint arXiv:2310.04988, 2023.
  67. 67.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. COMET: A neural framework for MT evaluation. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 2685–2702, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.213. URL https://aclanthology.org/2020.emnlp-main.213.
  68. 68.Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. COMET-22: Unbabel-IST 2022 submission for the metrics shared task. In Philipp Koehn, Loïc Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 578–585, Abu Dhabi, United Arab Emirates (Hybrid), December 2022a. Association for Computational Linguistics. URL https://aclanthology.org/2022.wmt-1.52.
  69. 69.Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task. In Philipp Koehn, Loïc Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 634–645, Abu Dhabi, United Arab Emirates (Hybrid), December 2022b. Association for Computational Linguistics. URL https://aclanthology.org/2022.wmt-1.60.
  70. 70.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  71. 71.Arleen Salles, Kathinka Evers, and Michele Farisco. Anthropomorphism in ai. AJOB neuroscience, 11(2):88–95, 2020.
  72. 72.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. QuestEval: Summarization asks for fact-based evaluation. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6594–6604, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.529. URL https://aclanthology.org/2021.emnlp-main.529.
  73. 73.Thibault Sellam, Dipanjan Das, and Ankur Parikh. BLEURT: Learning robust metrics for text generation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7881–7892, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.704. URL https://aclanthology.org/2020.acl-main.704.
  74. 74.Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. Personality traits in large language models. arXiv preprint arXiv:2307.00184, 2023.
  75. 75.Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov. Constructing benchmarks and interventions for combating hallucinations in llms, 2024.
  76. 76.Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 3607–3625, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.220. URL https://aclanthology.org/2023.emnlp-main.220.
  77. 77.Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. On early detection of hallucinations in factual question answering, 2023. URL https://arxiv.org/abs/2312.14183.
  78. 78.Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. Benchmarking hallucination in large language models based on unanswerable math word problem. CoRR, 2024.
  79. 79.Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona. Confidence improves self-consistency in llms. arXiv preprint arXiv:2502.06233, 2025.
  80. 80.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. Fine-tuning language models for factuality. arXiv preprint arXiv:2311.08401, 2023a.
  81. 81.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975, 2023b.
  82. 82.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  83. 83.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation, 2023.
  84. 84.Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson. " confidently nonsensical?": A critical survey on the perspectives and challenges of'hallucinations' in nlp. arXiv preprint arXiv:2404.07461, 2024.
  85. 85.Chaojun Wang and Rico Sennrich. On exposure bias, hallucination and domain shift in neural machine translation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 3544–3552, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.326. URL https://aclanthology.org/2020.acl-main.326.
  86. 86.Quanbin Wang and Ying Tan. Grammatical error detection with self attention by pairwise training. In 2020 International Joint Conference on Neural Networks, IJCNN 2020, Glasgow, United Kingdom, July 19-24, 2020, pp. 1–7. IEEE, 2020. doi: 10.1109/IJCNN48605.2020.9206715. URL https://doi.org/10.1109/IJCNN48605.2020.9206715.
  87. 87.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1112–1122. Association for Computational Linguistics, 2018. URL http://aclweb.org/anthology/N18-1101.
  88. 88.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2369–2380, 2018.
  89. 89.Fan Yin, Jayanth Srinivasa, and Kai-Wei Chang. Characterizing truthfulness in large language model generations with local intrinsic dimension. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, 2024.
  90. 90.Gal Yona, Roee Aharoni, and Mor Geva. Can large language models faithfully express their intrinsic uncertainty in words?, 2024. URL https://arxiv.org/abs/2405.16908.
  91. 91.Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. Attention satisfies: A constraint-satisfaction lens on factual errors of language models. In The Twelfth International Conference on Learning Representations, 2023.
  92. 92.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. ERNIE: Enhanced language representation with informative entities. In Anna Korhonen, David Traum, and Lluís Màrquez (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1441–1451, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1139. URL https://aclanthology.org/P19-1139.
  93. 93.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876, 2018.
  94. 94.Lina Zhou, Yongmei Shi, Jinjuan Feng, and Andrew Sears. Data mining for detecting errors in dictation speech recognition. IEEE Trans. Speech Audio Process., 13(5-1):681–688, 2005. doi: 10.1109/TSA.2005.851874. URL https://doi.org/10.1109/TSA.2005.851874.
  95. 95.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Representation engineering: A top-down approach to ai transparency, 2023. URL https://arxiv.org/abs/2310.01405.

Citation

MLA
Orgad, H., et al. “LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations”. arXiv, 2024, http://arxiv.org/abs/2410.02707v4.
APA
Orgad, H., Toker, M., Gekhman, Z., Reichart, R., Szpektor, I., Kotek, H., & Belinkov, Y. (2024). LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. arXiv. http://arxiv.org/abs/2410.02707v4
Chicago
Orgad, H., M. Toker, Z. Gekhman, et al. 2024. “LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations”. arXiv. http://arxiv.org/abs/2410.02707v4.
Harvard
Orgad, H. et al. (2024) “LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.02707v4.
Vancouver
1. Orgad H, Toker M, Gekhman Z, Reichart R, Szpektor I, Kotek H, Belinkov Y (2024) LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. arXiv

BibTeX

@article{orgad2024llms,
  title = {LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations},
  author = {Orgad, Hadas and Toker, Michael and Gekhman, Zorik and Reichart, Roi and Szpektor, Idan and Kotek, Hadas and Belinkov, Yonatan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.02707v4},
  eprint = {2410.02707}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors