In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
Shiqi ChenMiao XiongJunteng LiuZhengxuan WuTeng XiaoSiyang GaoJunxian He
Reveals that sharper in-context hidden state activations correlate with factual correctness in large language models and introduces Activation Decoding, an entropy-constrained generation method that significantly reduces hallucinations across standard question-answering benchmarks.
Large language models frequently generate factually incorrect outputs, known as hallucinations. This lack of reliability poses severe operational and reputational risks for organizations deploying generative artificial intelligence across decision-making, customer service, and knowledge management systems. While standard mitigation approaches rely on external knowledge retrieval or compute-heavy fine-tuning, the underlying mechanisms that cause models to produce false statements remain poorly understood.
The article investigates whether a model's internal representations—specifically the hidden activation states within intermediate layers—contain reliable signals that indicate factual correctness. Based on these insights, the authors aim to introduce and evaluate an unsupervised, inference-time decoding approach that suppresses hallucinations without requiring external knowledge bases or model retraining.
The authors analyzed token activations across intermediate transformer layers using factual question-answering benchmarks, including COUNTERFACT, TruthfulQA, TriviaQA, HotpotQA, and Natural Questions. They evaluated leading open-source models (the LLaMA-2-chat family across 7B, 13B, and 70B parameter sizes, alongside base LLaMA-2 and Mistral-7B models) and compared their proposed method against standard greedy decoding and prior controlled generation techniques such as DoLa and Inference-time Intervention.
The analysis produced several critical findings. First, factually correct predictions exhibit significantly sharper, more concentrated activation patterns across prompt tokens within deeper intermediate layers (such as layers 26 to 30), whereas incorrect outputs show diffuse, delayed activations. Second, measuring this sharpness via an entropy metric reliably distinguishes true answers from false ones, achieving an area under the receiver operating characteristic curve (AUROC) above 0.75. Third, incorporating this entropy metric into a constrained generation strategy—termed Activation Decoding—substantially enhances factuality across all model sizes. On TruthfulQA, the method improved the combined Truth*Info score by up to 8.6 points, and raised factual accuracy (F1 score) by up to 4.8 points on TriviaQA and 4.7 points on HotpotQA, with performance gains scaling positively with model size. Fourth, the method achieved a 7.3% reduction in inference latency compared to contrasting layer baselines by pre-computing prompt entropies, adding only a 23.4% overhead over standard greedy decoding.
These findings indicate that language models frequently possess correct factual knowledge within their internal representations even when default generation algorithms fail to elicit it. Deploying internal representation-based decoding enables organizations to improve factual precision and reduce evasive responses (such as 'I have no comment') at minimal operational cost, without modifying underlying model weights or building expensive retrieval pipelines.
Engineering and product teams deploying language models in production should evaluate Activation Decoding as a lightweight, plug-and-play intervention to enhance factuality. Where appropriate, technical teams can combine Activation Decoding with complementary layer-contrasting techniques (such as DoLa) to achieve compound accuracy gains.
Confidence in these findings is high across the tested open-ended and multiple-choice benchmarks. However, decision-makers should note key boundaries: this method only mitigates internal model elicitation failures and cannot correct factual errors caused by biased pre-training data, missing information, or outdated facts requiring external knowledge retrieval.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). Read this earlier inference-time intervention first to understand the truth-related activation steering that the source compares with its entropy-guided decoding method.
- Paper: On Large Language Models' Hallucination with Regard to Known Facts, Che Jiang et al. (2024). Its layer-by-layer account of known-fact hallucinations provides the mechanistic backdrop for the source’s analysis of how internal activations signal correct versus incorrect generation.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). This later work extends internal-state signals from guiding factual generation to predicting a model’s failures, building on the source’s premise that hidden activations reveal output correctness.
