Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Kevin LiuStephen CasperDylan Hadfield-MenellJacob Andreas

article2023EMNLP76 citations

Explains why internal probes often outperform direct language model outputs by categorizing query–probe disagreements into confabulation, deception, and heterogeneity, showing that superior probe accuracy stems primarily from better uncertainty calibration rather than intentional deception.

Listen

Artificial intelligence language models frequently generate factual errors, which creates risks for organizations relying on them as automated knowledge sources. Prior studies revealed that examining a model's internal representations via linear classifiers—known as probes—often yields more accurate factual judgments than asking the model directly through queries. This disparity has led to claims that language models exhibit deception by internally knowing the truth while outputting false statements.

The article investigates why model outputs disagree with their internal representations of truthfulness. It evaluates whether this mismatch stems from deliberate deception-like dynamics or alternative mechanisms, including differences in uncertainty calibration and performance across distinct subsets of data.

The researchers conducted an empirical evaluation using two autoregressive language models, GPT2-XL and GPT-J, across three question-answering benchmarks: BoolQ, SciQ, and CREAK. They compared direct query outputs against linear probes trained on the models' final internal hidden states. The study established a taxonomy classifying disagreements into three main types: model confabulation (where queries answer confidently despite probe uncertainty), deception (where both query and probe are confident but reach opposite conclusions), and heterogeneity (where probes and queries succeed on different subsets of inputs).

The analysis produced three primary findings. First, apparent deception is rare; high-confidence contradictions between probes and queries occurred infrequently across the benchmarks, showing up meaningfully only on the CREAK dataset. Second, the superior performance of probes is largely driven by better statistical calibration on uncertain answers rather than a greater volume of confident, correct facts. Third, queries and probes exhibit complementary strengths on different subsets of data: by combining both approaches into a weighted ensemble, the researchers increased overall accuracy beyond either individual method in four out of six experimental settings, such as raising GPT-J accuracy on SciQ from 84.1% via query and 88.6% via probe to 90.5% via the ensemble.

These findings indicate that internal-versus-external discrepancies stem from differing computational prediction pathways rather than an underlying intent to mislead. For organizational leadership, this clarifies that anthropomorphic framing around model lying is largely inaccurate. However, it also highlights that internal probing cannot fully resolve language model unreliability, as models still make significant factual errors across both pathways.

Organizations developing or deploying factual language systems should consider ensembling output probabilities with internal probe signals to achieve incremental accuracy gains and better uncertainty estimates. When high accuracy is critical, decision-makers should fine-tune underlying models directly on domain data rather than relying solely on probing pretrained models, as fine-tuning consistently outperformed zero-shot queries and probes on datasets like SciQ and CREAK.

Confidence in these findings is supported by consistent error distributions across different model sizes and probe regularizations. Nonetheless, decision-makers should exercise caution: the study was limited to binary question answering, evaluated a single prompt template per task, and tested models that continue to generate substantial factual errors, making them unsuitable for fully autonomous deployment in mission-critical settings without human oversight.

  • Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). Its Contrast-Consistent Search method establishes how hidden-state probes can recover factual knowledge that ordinary language-model answers fail to express, the central premise this study tests.
  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its analysis of models’ self-evaluation and calibration provides essential context for this study’s finding that probe advantages often reflect better handling of uncertainty.
  • Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Its TruthfulQA benchmark frames the problem of fluent, confident falsehoods that motivates studying gaps between internal truth signals and generated answers.
  • Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Its methods for eliciting factual knowledge from language models establish the earlier problem of separating what models know from what their prompts reveal.
Cover for Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Abstract

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these two procedures sometimes disagree, and that probes tend to be more accurate than LM outputs. This has led some researchers to conclude that LMs “lie” or otherwise encode non-cooperative communicative intents. Is this an accurate description of today’s LMs, or can query–probe disagreement arise in other ways? We identify three different classes of disagreement, which we term confabulation, deception, and heterogeneity. In many cases, the superiority of probes is simply attributable to better calibration on uncertain answers rather than a greater fraction of correct, high-confidence answers. In some cases, queries and probes perform better on different subsets of inputs, and accuracy can further be improved by ensembling the two.¹

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 The success of probes over queries can largely be explained by better calibration
  • 4 Most disagreements result from confabulation and heterogeneity
  • 5 Queries and probes are complementary
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethical Considerations
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Three operational forms of query–probe disagreement

    definition

    For a language model, a query ranks candidate answers from the probabilities it assigns to them, while a probe estimates whether each question–answer pair is correct from the model’s hidden representation. The paper distinguishes three reasons these predictions can disagree:

    • Model confabulation: the query gives an incorrect answer with high confidence while the probe has low confidence. In this case, the probe’s advantage may come from being better calibrated to uncertainty, rather than from producing more correct high-confidence answers.
    • “Deception”: the probe is confidently correct while the query is confidently incorrect. This is an operational name for a prediction pattern, not evidence that a model has communicative intent or believes one thing while choosing to say another.
    • Heterogeneity: query and probe accuracy differs across subsets of inputs, so each can be more effective on examples where the other is less effective.

    These behaviors can coexist within a model and are distinct from agreement and probe errors.

  2. Knowl 2 — Evaluation and probe-training protocol

    experimental setup

    The study compared GPT2-XL and GPT-J on BoolQ, SciQ, and CREAK as binary question-answering tasks. BoolQ and CREAK answers were represented by the strings true and false, with language-model output probabilities renormalized over those two choices. For SciQ, each question was evaluated using its correct answer and one selected distractor.

    For querying, candidate answers were ranked using the language model’s answer probabilities. For probing, a logistic-regression linear classifier was trained on the dataset’s training split to classify question–answer pairs as correct or incorrect. Its input was the model’s final-layer, final hidden state for the pair; the classifier’s correctness scores were normalized across candidate answers to form an answer distribution. Accuracy was evaluated on held-out test data.

  3. Knowl 3 — Query and probe accuracy across models and datasets

    data/table

    The table reports test accuracy (percent) for direct language-model queries and trained linear probes. Probes are more accurate in all six model–dataset comparisons; the gap is especially large on CREAK, while the GPT-J BoolQ difference is small.

    BoolQ SciQ CREAK
    Method GPT2-XL GPT-J GPT2-XL GPT-J GPT2-XL GPT-J
    Query 61.8 61.8 76.1 84.1 51.1 50.4
    Probe 62.9 62.5 78.8 88.6 63.5 71.0
  4. Knowl 4 — Probe accuracy advantage is associated with better calibration

    empirical result

    In the study’s calibration analysis of GPT-J, probe confidence tracked empirical correctness substantially better than query confidence across BoolQ, SciQ, and CREAK. Query predictions were reasonably calibrated on SciQ but poorly calibrated on BoolQ and CREAK. The authors therefore argue that probes’ higher accuracy can often be explained by better calibration of uncertainty, rather than by probes simply producing a larger share of correct, high-confidence predictions.

  5. Knowl 5 — Most disagreements are not confident query errors against correct probes

    empirical result

    Across GPT2-XL and GPT-J, the joint distribution of query and probe predictions varied substantially by dataset. On BoolQ and SciQ, cases where the query was correct and the probe incorrect occurred nearly as often as cases with the reverse outcome; highly confident contradictory predictions were rare. The study found a notable number of “deception”-pattern cases only on CREAK, where probe errors were also common. Overall, most mismatches—including on SciQ, where probes were much more accurate—did not fit the confident-probe-correct/confident-query-wrong pattern.

  6. Knowl 6 — A validation-tuned query–probe ensemble can improve test accuracy

    model/method

    The study combined the query and probe answer distributions by a convex mixture:

    pensemble(a∣q)=λpprobe(a∣q)+(1−λ)pquery(a∣q),p_{\mathrm{ensemble}}(a\mid q)=\lambda p_{\mathrm{probe}}(a\mid q)+(1-\lambda)p_{\mathrm{query}}(a\mid q),

    where qq is a question, aa is a candidate answer, each component is a probability distribution over candidate answers, and λ∈[0,1]\lambda\in[0,1] is selected using 500 validation examples from the relevant dataset. The mixture was then evaluated on the test split. It improved over the probe in four of six model–dataset combinations: GPT-J on BoolQ (62.7% versus 62.5% probe accuracy), GPT2-XL on SciQ (79.7% versus 78.8%), GPT-J on SciQ (90.5% versus 88.6%), and GPT-J on CREAK (71.3% versus 71.0%). No ensemble accuracy was reported for GPT2-XL on BoolQ or CREAK. The gains show that query and probe predictions can provide complementary information.

  7. Knowl 7 — Comparison with fine-tuned GPT2-XL queries

    empirical result

    As a comparison with the many-example probe, the authors fine-tuned GPT2-XL on correct question–answer pairs and evaluated its queries. Probe predictions had 1.7% better accuracy than fine-tuned queries on BoolQ; fine-tuned queries had 6.9% better accuracy on SciQ and 1.9% better accuracy on CREAK. Thus, the probe was not uniformly more accurate than a fine-tuned language model across datasets.

  8. Knowl 8 — Sparse-probe results are mostly stable except at extreme sparsity

    empirical result

    The authors also trained linear probes with an ℓ1\ell_1 penalty varied over {0,0.01,0.03,0.1}\{0, 0.01, 0.03, 0.1\} to encourage sparse solutions, and assessed probe accuracy, sparsity, and the distribution of disagreement types. They report that, except at extremely high sparsity, accuracy and disagreement distributions remained similar to those from the main experiments. This robustness check concerns sparse versions of the same linear probing approach; it does not establish that linear probes capture all relevant information in the model representations.

  9. Knowl 9 — Dataset, prompt, and probe choices limit generalization

    limitation

    The distribution of disagreement types differed substantially across BoolQ, SciQ, and CREAK, so the paper cautions that its dataset-specific findings may not predict results on future datasets. Each experiment used a single prompt template, and alternative prompts—particularly for CREAK—could change the disagreement distribution. The main experiments used linear probes only; the authors infer from the benefits of combining queries and probes that some useful information is not captured by the linear probe. The study therefore does not establish that its observed disagreement patterns generalize across prompts, datasets, or probe classes.

Coverage note — No other substantial contributed result was omitted; background, related work, and ethical discussion were excluded because they do not add a distinct method or empirical finding to reconstruct the paper’s main contribution.

References

  1. 1.Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, and Zeerak Talat. 2023. Mirages: On anthropomorphism in dialogue systems. arXiv preprint arXiv:2305.09800.
  2. 2.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  3. 3.Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when it’s lying. arXiv preprint arXiv:2304.13734.
  4. 4.Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.
  5. 5.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations.
  6. 6.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the Conference of the North American Chapter of the Association for Comptuational Linguistics.
  7. 7.Benj Edwards. 2023. Why ChatGPT and Bing Chat are so good at making things up. Ars Technica.
  8. 8.Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175.
  9. 9.Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders. 2021. Truthful AI: Developing and governing AI that does not lie. arXiv preprint arXiv:2110.06674.
  10. 10.Evan Hernandez and Jacob Andreas. 2021. The low-dimensional linear geometry of contextualized word representations. In Proceedings of the Conference on Natural Language Learning.
  11. 11.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  12. 12.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  13. 13.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the Annual Meeting of the Association for Computational Linguistics.
  14. 14.James Edwin Mahon. 2008. The definition of lying and deception.
  15. 15.Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824.
  16. 16.Marianna Martindale, Marine Carpuat, Kevin Duh, and Paul McNamee. 2019. Identifying fluently inadequate output in neural and statistical machine translation. In Proceedings of Machine Translation Summit XVII: Research Track.
  17. 17.Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872.
  18. 18.Yasumasa Onoe, Michael JQ Zhang, Eunsol Choi, and Greg Durrett. 2021. CREAK: A dataset for commonsense reasoning over entity knowledge. In NeurIPS Datasets And Benchmarks.
  19. 19.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the Conference on Empirical Methods in Natural Language Processing.
  20. 20.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  21. 21.Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell. 2022. Linear adversarial concept erasure. In Proceedings of the International Conference on Machine Learning.
  22. 22.Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2020. Probing the probing paradigm: Does probing accuracy entail task relevance? In Proceedings of the Annual Meeting of the European Association for Computational Linguistics.
  23. 23.Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022. Neural theory-of-mind? On the limits of social intelligence in large lms. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  24. 24.Murray Shanahan. 2022. Talking about large language models. arXiv preprint arXiv:2212.03551.
  25. 25.Elena Voita and Ivan Titov. 2020. Information-theoretic probing with minimum description length. In Proceedings of the Conference on Empirical Methods in Natural Language Processing.
  26. 26.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  27. 27.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023. Representation engineering: A top-down approach to AI transparency. arXiv preprint arXiv:2310.01405.

Citation

MLA
Liu, K., et al. “Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4791–97, https://doi.org/10.18653/v1/2023.emnlp-main.291.
APA
Liu, K., Casper, S., Hadfield-Menell, D., & Andreas, J. (2023). Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4791–4797. https://doi.org/10.18653/v1/2023.emnlp-main.291
Chicago
Liu, K., S. Casper, D. Hadfield-Menell, and J. Andreas. 2023. “Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4791–97. https://doi.org/10.18653/v1/2023.emnlp-main.291.
Harvard
Liu, K. et al. (2023) “Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4791–4797. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.291.
Vancouver
1. Liu K, Casper S, Hadfield-Menell D, Andreas J (2023) Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4791–4797

BibTeX

@inproceedings{liu-etal-2023-cognitive,
    title = "Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?",
    author = "Liu, Kevin  and
      Casper, Stephen  and
      Hadfield-Menell, Dylan  and
      Andreas, Jacob",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.291/",
    doi = "10.18653/v1/2023.emnlp-main.291",
    pages = "4791--4797"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/