In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation

Shiqi ChenMiao XiongJunteng LiuZhengxuan WuTeng XiaoSiyang GaoJunxian He

article2024ICML58 citations

Reveals that sharper in-context hidden state activations correlate with factual correctness in large language models and introduces Activation Decoding, an entropy-constrained generation method that significantly reduces hallucinations across standard question-answering benchmarks.

Listen

Large language models frequently generate factually incorrect outputs, known as hallucinations. This lack of reliability poses severe operational and reputational risks for organizations deploying generative artificial intelligence across decision-making, customer service, and knowledge management systems. While standard mitigation approaches rely on external knowledge retrieval or compute-heavy fine-tuning, the underlying mechanisms that cause models to produce false statements remain poorly understood.

The article investigates whether a model's internal representations—specifically the hidden activation states within intermediate layers—contain reliable signals that indicate factual correctness. Based on these insights, the authors aim to introduce and evaluate an unsupervised, inference-time decoding approach that suppresses hallucinations without requiring external knowledge bases or model retraining.

The authors analyzed token activations across intermediate transformer layers using factual question-answering benchmarks, including COUNTERFACT, TruthfulQA, TriviaQA, HotpotQA, and Natural Questions. They evaluated leading open-source models (the LLaMA-2-chat family across 7B, 13B, and 70B parameter sizes, alongside base LLaMA-2 and Mistral-7B models) and compared their proposed method against standard greedy decoding and prior controlled generation techniques such as DoLa and Inference-time Intervention.

The analysis produced several critical findings. First, factually correct predictions exhibit significantly sharper, more concentrated activation patterns across prompt tokens within deeper intermediate layers (such as layers 26 to 30), whereas incorrect outputs show diffuse, delayed activations. Second, measuring this sharpness via an entropy metric reliably distinguishes true answers from false ones, achieving an area under the receiver operating characteristic curve (AUROC) above 0.75. Third, incorporating this entropy metric into a constrained generation strategy—termed Activation Decoding—substantially enhances factuality across all model sizes. On TruthfulQA, the method improved the combined Truth*Info score by up to 8.6 points, and raised factual accuracy (F1 score) by up to 4.8 points on TriviaQA and 4.7 points on HotpotQA, with performance gains scaling positively with model size. Fourth, the method achieved a 7.3% reduction in inference latency compared to contrasting layer baselines by pre-computing prompt entropies, adding only a 23.4% overhead over standard greedy decoding.

These findings indicate that language models frequently possess correct factual knowledge within their internal representations even when default generation algorithms fail to elicit it. Deploying internal representation-based decoding enables organizations to improve factual precision and reduce evasive responses (such as 'I have no comment') at minimal operational cost, without modifying underlying model weights or building expensive retrieval pipelines.

Engineering and product teams deploying language models in production should evaluate Activation Decoding as a lightweight, plug-and-play intervention to enhance factuality. Where appropriate, technical teams can combine Activation Decoding with complementary layer-contrasting techniques (such as DoLa) to achieve compound accuracy gains.

Confidence in these findings is high across the tested open-ended and multiple-choice benchmarks. However, decision-makers should note key boundaries: this method only mitigates internal model elicitation failures and cannot correct factual errors caused by biased pre-training data, missing information, or outdated facts requiring external knowledge retrieval.

arXiv: 2403.01548hkust-nlp/Activation_Decoding
Cover for In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation

Abstract

Large language models (LLMs) frequently hallucinate and produce factual errors, yet our understanding of why they make these errors remains limited. In this study, we delve into the underlying mechanisms of LLM hallucinations from the perspective of inner representations, and discover a salient pattern associated with hallucinations: correct generations tend to have sharper context activations in the hidden states of the in-context tokens, compared to the incorrect ones. Leveraging this insight, we propose an entropy-based metric to quantify the “sharpness” among the in-context hidden states and incorporate it into the decoding process to formulate a constrained decoding approach. Experiments on various knowledge-seeking and hallucination benchmarks demonstrate our approach’s consistent effectiveness, for example, achieving up to an 8.6 point improvement on TruthfulQA. We believe this study can improve our understanding of hallucinations and serve as a practical solution for hallucination mitigation. Code is publicly available at https://github.com/hkust-nlp/Activation_Decoding.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Diving into Internal Representations
  • 3.1 Notation
  • 3.2 Experimental Setup for the Case Study
  • 3.3 Finding 1: Activation implies answer correctness
  • 3.4 Finding 2: The contextual entropy of correct answers is consistently smaller than incorrect ones.
  • 3.5 Finding 3: Contextual entropy can calibrate the next token prediction
  • 4 Activation Decoding
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Results
  • 5.3 Qualitative Study: What types of errors can our method address?
  • 6 Conclusions and Discussion
  • References
  • A Model Generalization
  • B Dataset Curation
  • C Hyperparameter Generalization
  • D Inference Efficiency

Knowls

  1. Knowl 1 — Subject-token activation is associated with answer correctness

    empirical result

    In a LLaMA-2-chat-7B study on COUNTERFACT, a generated answer token was counted as activated when it ranked among the top 50 vocabulary tokens projected from the last subject token’s hidden state at layer 26 of 32. Correct answers were activated more often than incorrect answers: Raw-CFT had rates of 81.29% (226 activated, 52 unactivated) versus 24.14% (21 activated, 66 unactivated); GF-CFT had rates of 63.00% (441 activated, 259 unactivated) versus 36.92% (120 activated, 205 unactivated). Raw-CFT’s incorrect answers were model outputs manually judged by the authors, whereas GF-CFT’s were the dataset’s constructed false answers. The result supports an association between a candidate answer being represented at the subject position and its factual correctness; activation is not itself a guarantee of correctness.

  2. Knowl 2 — Contextual entropy quantifies in-context activation sharpness

    definition

    Let a prompt contain tokens C=(v1,…,vt)C=(v_1,\ldots,v_t), and let vv be a candidate next token. At a selected transformer layer ll, let xi(l)x_i^{(l)} be the hidden-state vector at prompt position ii, and let WW be the language-model output head mapping hidden states to vocabulary logits. Define the activation score as s(i,v)=softmax⁡(Wxi(l))vs(i,v)=\operatorname{softmax}(W x_i^{(l)})_v, where the softmax is over the vocabulary. Normalize these scores across the tt prompt positions as p~i(v∣C)=exp⁡(s(i,v))/∑m=1texp⁡(s(m,v))\tilde p_i(v\mid C)=\exp(s(i,v))/\sum_{m=1}^{t}\exp(s(m,v)). The contextual entropy is E(v,C)=−∑i=1tp~i(v∣C)log⁡p~i(v∣C)E(v,C)=-\sum_{i=1}^{t}\tilde p_i(v\mid C)\log \tilde p_i(v\mid C). A lower value means the candidate’s activation is more concentrated on a smaller number of prompt positions, which the paper treats as greater in-context sharpness and an indicator—not a guarantee—of factual correctness.

  3. Knowl 3 — Activation Decoding favors candidates with sharper prompt activations

    model/method

    Activation Decoding adjusts an autoregressive language model’s next-token distribution using contextual entropy. For candidate token vv, let q(v)q(v) be its original next-token probability and let E(v,C)E(v,C) be its entropy with respect to the fixed input prompt CC. The adjusted probability is proportional to q(v)exp⁡(−λE(v,C))q(v)\exp(-\lambda E(v,C)), where λ∈[0,1]\lambda\in[0,1] controls the intervention strength; thus lower-entropy candidates receive a relative boost. The method computes entropy from the prompt tokens only, excluding generated tokens, so the entropy values remain fixed during continuation. At each decoding step, it first selects tokens with original probability at least 0.10.1 times the maximum next-token probability, adjusts those candidates, and then applies the chosen decoding procedure (including greedy decoding or beam search) to the adjusted distribution. For efficiency, it can compute and cache entropy for every vocabulary token from the prompt’s layer-ll hidden states before generation; the paper describes a 32,000-entry entropy vector for LLaMA-2.

  4. Knowl 4 — Contextual entropy separates ground-truth and false answers

    empirical result

    On GF-CFT, the contextual entropy of ground-truth answers was generally lower than that of constructed false answers, consistent with correct candidates having sharper activation over prompt positions. Using hidden states from LLaMA-2-chat-7B at layer 26 yielded an AUROC of 0.75 for distinguishing the two groups; layer 28 yielded an AUROC of 0.76. These results establish entropy as a useful factual-error signal in this dataset and setup, rather than a universally reliable correctness test.

  5. Knowl 5 — Entropy improves answer-level factual-error detection

    empirical result

    For answer-level correctness classification on COUNTERFACT-derived Raw-CFT and GF-CFT, the authors compared AUROC for self-evaluation, sequence logits, DoLa-adjusted logits, subject-token activation, and sequence logits combined with contextual entropy. With the entropy computed from layer 27 of LLaMA-2-chat-7B, the reported AUROC points were: GF-CFT—self-eval 66.83, logits 70.79, DoLa 68.96, subject activation 71.59, and logits plus entropy 72.65; Raw-CFT—self-eval 61.46, logits 72.05, DoLa 72.27, subject activation 73.21, and logits plus entropy 74.15. Logits plus entropy performed best among these alternatives on both datasets and exceeded the logits baseline by 1.86 and 2.10 AUROC points, respectively.

  6. Knowl 6 — TruthfulQA factuality–informativeness balance improves

    empirical result

    On open-ended TruthfulQA, using LLaMA-2-chat models and scores reported on a 0–100 scale, Activation Decoding increased the TruthInfo score over greedy decoding for all three model sizes. For 7B, greedy versus Activation Decoding scored 55.8 versus 59.1 on TruthInfo, with Truth 62.9 versus 63.2, Info 92.8 versus 95.8, and Reject 12.7 versus 9.7. For 13B, the respective scores were 57.5 versus 62.3 on TruthInfo, 66.5 versus 64.3 on Truth, 91.1 versus 98.0 on Info, and 13.6 versus 5.5 on Reject. For 70B, they were 47.1 versus 55.7 on TruthInfo, 68.8 versus 65.7 on Truth, 78.3 versus 90.0 on Info, and 30.0 versus 15.7 on Reject. Thus the joint score improved by 3.3, 4.8, and 8.6 points, respectively, alongside substantially higher informativeness; Truth scores declined for the 13B and 70B models.

  7. Knowl 7 — Knowledge-seeking QA F1 improves across model sizes

    empirical result

    On open-ended TriviaQA, HotpotQA, and Natural Questions (NQ), Activation Decoding improved F1 over greedy decoding for every LLaMA-2-chat size reported. F1 scores, listed as greedy → Activation Decoding, were: 7B—TriviaQA 44.3 → 46.4, HotpotQA 20.1 → 21.1, NQ 20.4 → 21.4; 13B—TriviaQA 60.9 → 62.8, HotpotQA 21.7 → 26.4, NQ 28.9 → 32.5; 70B—TriviaQA 68.4 → 73.2, HotpotQA 25.5 → 30.1, NQ 34.1 → 37.8. The gains reached 4.8 points on TriviaQA, 4.7 on HotpotQA, and 3.7 on NQ. These are benchmark results for the paper’s selected hyperparameters, not a claim that every decoding configuration will yield the same gains.

  8. Knowl 8 — TruthfulQA-selected settings transfer to other QA datasets

    empirical result

    The authors tested out-of-domain hyperparameter selection by choosing the informative layer and entropy weight on TruthfulQA, then applying them without dataset-specific tuning to TriviaQA, HotpotQA, and NQ. In this setting, F1 scores for greedy decoding versus Activation Decoding were: LLaMA-2-chat-7B—TriviaQA 44.3 vs. 44.4, HotpotQA 20.1 vs. 20.8, NQ 20.4 vs. 21.0; 13B—60.9 vs. 62.7, 21.7 vs. 23.3, 28.9 vs. 32.4; 70B—68.4 vs. 73.2, 25.5 vs. 27.4, 34.1 vs. 37.4. Activation Decoding exceeded greedy decoding in all nine comparisons, showing transfer of the TruthfulQA-selected settings across these datasets and model sizes.

  9. Knowl 9 — Activation Decoding adds less latency than DoLa in the tested setup

    empirical result

    Inference time was measured on 722 Natural Questions examples using LLaMA-2-chat-7B on one NVIDIA Tesla A800 80GB GPU. Mean times were 167.93 seconds for greedy decoding, 222.38 seconds for DoLa, and 207.24 seconds for Activation Decoding. In this experiment, Activation Decoding was 7.3% faster than DoLa but 23.4% slower than greedy decoding. The paper attributes its relative speed advantage over DoLa to avoiding DoLa’s contrast-layer selection computation; the timing comparison is specific to this model, dataset sample, and hardware.

  10. Knowl 10 — The method depends on factual knowledge being present in the model

    limitation

    Activation Decoding is intended to elicit useful signals from a model’s existing internal representations, without external knowledge. Its underlying assumption is that relevant ground-truth knowledge is often already encoded in the hidden states of prompt tokens but is not reliably elicited during ordinary decoding. Consequently, it cannot supply missing knowledge or correct errors caused by outdated or erroneous training data. More generally, the paper notes that representation-based interventions do not provide a universal signal for every error type: effectiveness can vary by dataset, and correcting one error can introduce another, creating an inherent performance trade-off.

Coverage note — Supplementary Mistral and base-model experiments, TruthfulQA multiple-choice and Exact Match results, GPT-4 response-quality ratings, and qualitative examples are omitted because the knowls prioritize the central activation signal, decoding method, and main quantitative findings.

References

  1. 1.Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. Self-rag: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511, 2023. URL https://arxiv.org/abs/2310.11511.
  2. 2.Chen, S., Zhao, Y., Zhang, J., Chern, I.-C., Gao, S., Liu, P., and He, J. Felm: Benchmarking factuality evaluation of large language models. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023. URL http://arxiv.org/abs/2310.00741.
  3. 3.Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J., and He, P. Dola: Decoding by contrasting layers improves factuality in large language models. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/pdf/2309.03883.pdf.
  4. 4.Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. Knowledge neurons in pretrained transformers. In Association for Computational Linguistics (ACL), 2022. URL https://arxiv.org/abs/2104.08696.
  5. 5.Dar, G., Geva, M., Gupta, A., and Berant, J. Analyzing transformers in embedding space. In Association for Computational Linguistics (ACL), 2023. URL https://arxiv.org/abs/2209.02535.
  6. 6.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. A mathematical framework for transformer circuits. In Transformer Circuits Thread, 2021. URL https://transformer-circuits.pub/2021/framework/index.html.
  7. 7.Geva, M., Schuster, R., Berant, J., and Levy, O. Transformer feed-forward layers are key-value memories. In Association for Computational Linguistics (ACL), 2022. URL https://arxiv.org/abs/2012.14913.
  8. 8.Geva, M., Bastings, J., Filippova, K., and Globerson, A. Dissecting recall of factual associations in autoregressive language models. In Empirical Methods in Natural Language Processing (EMNLP), 2023. URL https://arxiv.org/abs/2304.14767.
  9. 9.Halawi, D., Denain, J.-S., and Steinhardt, J. Overthinking the truth: Understanding how language models process false demonstrations. arXiv preprint arXiv:2307.09476, 2023. URL https://arxiv.org/abs/2307.09476.
  10. 10.Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232, 2023. URL https://arxiv.org/abs/2311.05232.
  11. 11.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. Survey of hallucination in natural language generation. In ACM Computing Surveys, 2023. URL https://arxiv.org/abs/2202.03629.
  12. 12.Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G. Active retrieval augmented generation. In Empirical Methods in Natural Language Processing (EMNLP), 2023. URL https://arxiv.org/abs/2305.06983.
  13. 13.Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Association for Computational Linguistics (ACL), 2017. URL https://arxiv.org/abs/1705.03551.
  14. 14.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. Language models (mostly) know what they know. In Findings of Association for Computational Linguistics (ACL), 2023. URL https://arxiv.org/abs/2207.05221.
  15. 15.Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., and McHardy, R. Challenges and applications of large language models. In arXiv preprint arXiv:2307.10169, 2023. URL https://arxiv.org/abs/2307.10169.
  16. 16.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. Natural questions: A benchmark for question answering research. In Transactions of the Association of Computational Linguistics (TACL), 2019. URL https://aclanthology.org/Q19-1026/.
  17. 17.Li, K., Patel, O., Viegas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023a. URL https://arxiv.org/abs/2306.03341.
  18. 18.Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M. Contrastive decoding: Open-ended text generation as optimization. In Association for Computational Linguistics (ACL), 2023b. URL https://arxiv.org/abs/2210.15097.
  19. 19.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Association for Computational Linguistics (ACL), 2022. URL https://arxiv.org/abs/2109.07958.
  20. 20.Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Locating and editing factual associations in gpt. In Advances in Neural Information Processing Systems, 2022.
  21. 21.Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. Progress measures for grokking via mechanistic interpretability. In International Conference on Learning Representations (ICLR), 2023. URL https://arxiv.org/abs/2301.05217.
  22. 22.Olah, C. Mechanistic interpretability, variables, and the importance of interpretable bases. In Transformer Circuits Thread, 2022. URL URLhttps://transformer-circuits.pub/2022/mech-interp-essay/index.html.
  23. 23.OpenAI. Introducing chatgpt. URL https://openai.com/blog/chatgpt, 2022.
  24. 24.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023. URL https://arxiv.org/abs/2303.08774.
  25. 25.Pan, L., Saxon, M., Xu, W., Nathani, D., Wang, X., and Wang, W. Y. Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies. In arXiv preprint arXiv:2308.03188, 2023. URL https://arxiv.org/abs/2308.03188.
  26. 26.Ram, O., Bezalel, L., Zicher, A., Belinkov, Y., Berant, J., and Globerson, A. What are you token about? dense retrieval as distributions over the vocabulary. In Association for Computational Linguistics (ACL), 2023a. URL https://arxiv.org/abs/2212.10380.
  27. 27.Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y. In-context retrieval-augmented language models. In Association for Computational Linguistics (ACL), 2023b. URL https://arxiv.org/abs/2302.00083.
  28. 28.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. In arXiv preprint arXiv:2307.09288, 2023. URL https://arxiv.org/abs/2307.09288.
  29. 29.Wang, C., Liu, X., Yue, Y., Tang, X., Zhang, T., Jiayang, C., Yao, Y., Gao, W., Hu, X., Qi, Z., et al. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity. In arXiv preprint arXiv:2310.07521, 2023. URL https://arxiv.org/abs/2310.07521.
  30. 30.Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. In International Conference on Learning Representations (ICLR), 2023. URL https://arxiv.org/abs/2306.13063.
  31. 31.Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Empirical Methods in Natural Language Processing (EMNLP), 2018. URL https://arxiv.org/abs/1809.09600.
  32. 32.Yu, W., Zhang, Z., Liang, Z., Jiang, M., and Sabharwal, A. Improving language models via plug-and-play retrieval feedback. In arXiv preprint arXiv:2305.14002, 2023. URL https://arxiv.org/abs/2305.14002.
  33. 33.Yuksekgonul, M., Chandrasekaran, V., Jones, E., Gunasekar, S., Naik, R., Palangi, H., Kamar, E., and Nushi, B. Attention satisfies: A constraint-satisfaction lens on factual errors of language models. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/abs/2309.15098.
  34. 34.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. In arXiv preprint arXiv:2310.01405, 2023. URL https://arxiv.org/abs/2310.01405.

Citation

MLA
Chen, S., et al. “In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation”. arXiv, 2024, https://doi.org/10.48550/arxiv.2403.01548.
APA
Chen, S., Xiong, M., Liu, J., Wu, Z., Xiao, T., Gao, S., & He, J. (2024). In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation. arXiv. https://doi.org/10.48550/arxiv.2403.01548
Chicago
Chen, S., M. Xiong, J. Liu, et al. 2024. “In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2403.01548.
Harvard
Chen, S. et al. (2024) “In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation”. arXiv. Available at: https://doi.org/10.48550/arxiv.2403.01548.
Vancouver
1. Chen S, Xiong M, Liu J, Wu Z, Xiao T, Gao S, He J (2024) In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation. https://doi.org/10.48550/arxiv.2403.01548

BibTeX

@misc{https://doi.org/10.48550/arxiv.2403.01548,
  doi = {10.48550/ARXIV.2403.01548},
  url = {https://arxiv.org/abs/2403.01548},
  author = {Chen, Shiqi and Xiong, Miao and Liu, Junteng and Wu, Zhengxuan and Xiao, Teng and Gao, Siyang and He, Junxian},
  keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/